How to Fix autoconf-style Config Probing

Posted on 8 Oct 2026 by Boris Kolpackov

Configuration probing as implemented in autoconf/CMake/etc involves compiling and linking a test program to determine whether a particular feature, such as a function, is available on the platform being targeted. For example, we may prefer to use the strl*() family of functions in our codebase. However, these functions are not (yet) standard and are not provided by all libc implementations. As a result, we may wish to detect whether they are present and if not, provide fallback implementations or use alternatives. One way to do this detection would be to compile and link a test program that tries to use the functions we are interested in. If that succeeds, then we conclude the functions are available.

On the face of it, this approach is appealing. In particular, it is adaptable in the sense that we don't have to do anything to support platforms that may not even exist yet. For example, if someone decides to write yet another libc for Linux, we don't have to do anything to support it – the existing strl*() probes will sort it out. In fact, even already released versions of our project will automagically support this new libc.

This approach does have a few annoying problems. Here are the main ones:

  1. It is wasteful: There is no need to keep compiling the strl*() probes on, say, FreeBSD, where these functions were available for eons. At the limit this becomes absurd, like keep probing for a feature while the latest target that doesn't have it would not even be able to perform the probe. For a good example, see A Generation Lost in the Bazaar.
  2. It is brittle: We decide that a feature is absent based on the failure to compile/link a test program. But a lot of other things can lead to a failure to compile or link: mistakes in the test, misconfigured build, missing feature test macros such as _GNU_SOURCE, etc.

    For example, a lot of weeping and gnashing of teeth was recently caused by false negatives due to sloppily written probes. They stopped compiling because GCC and Clang stopped accepting certain long-deprecated C constructs.

    The failure mode is also insidious: a false negative silently leads to the feature not being used, leading to missing functionality, suboptimal performance, etc.

  3. It is slow: While compiling a single probe doesn't take long, compiling several hundreds is noticeable. To exacerbate the problem, both autoconf and CMake do it serially.
  4. It lacks change-tracking: Existing tools (autoconf, CMake) do not re-run the relevant probes when their inputs change. For example, strl*() were added in glibc 2.38. If we upgraded from 2.37, we would want all the already configured projects on our machine to detect the change and start using the newly available functions.

Solving the first problem (wastefulness) requires a completely different approach. One alternative is to use what we can call "expectation-based configuration": we assume a feature is available if certain conditions are met. For example, for strl*() we could assume these functions are available if we are targeting FreeBSD or glibc version 2.38 or later (of course, a complete implementation would also need to check for other platforms and/or libc implementations). This approach has been successfully used in build2 on configuration-heavy projects such as Qt and FFmpeg (see libbuild2-autoconf for details).

Ok, let's say we still wish to do configuration probing for some reason or for some special cases. Can we solve, or at least mitigate, the remaining problems? Let's save the brittleness problem for last and take a stab at the remaining two: slowness and lack of change-tracking.

A high-level view of what we are doing during configuration probing can be summed up like this: we are compiling and linking a number of test programs, except that the result we are after is not the programs but rather the status: whether the compilation and linking succeeded or failed. We would like to do this in parallel and also keep track of changes to inputs: test source itself, recursive set of headers included by it, compile/link options, etc.

Doesn't the shape of this problem look familiar? What existing problem requires us to compile and link a bunch of source files in parallel and with proper change-tracking? That's right, this is how we build our software with existing build systems. Even make can do this reasonably well.

Apparently, CMake generates an individual project per each probe and then runs the underlying build system to build it. But it neither uses this to run multiple probes in parallel nor to track changes.

So couldn't we just use the build system to do the probing? And while at it couldn't we get rid of the whole separate configuration/project generation step?

It could work like this: we run the build system to update our project, it builds (or re-builds) the probes as necessary and then uses the resulting information to build our project source code. Specifically to our strl*() example, the build system would compile and link strlcpy.c and strlcat.c fallback implementation if the probes for these functions returned negative results.

Another advantage of using a proper build system for probes is the ability to establish dependencies between probe results. For example, there is no use wasting time probing for strl*() if there is no <string.h>.

There is one snag, though: the results of the probes need to be known when loading and evaluating the buildfiles. For example, we decide whether to include strlcpy.c and strlcat.c into the build while evaluating buildfile definitions. Here is a GNU make-based illustration:

hello: hello.o
hello.o: hello.c

ifndef have_strlcpy
  hello: strlcpy.o
  strlcpy.o: strlcpy.c
else
  CPPFLAGS += -DHAVE_STRLCPY
endif

ifndef have_strlcat
  hello: strlcat.o
  strlcat.o: strlcat.c
else
  CPPFLAGS += -DHAVE_STRLCAT
endif

In the above example, have_strlcpy and have_strlcpy would need to be known when make is evaluating this makefile but if running the corresponding probes is part of the overall build, then their values are only known later, once the makefile has been evaluated and make starts actually building the targets.

We could easily overcome this snag if we had the ability to pause loading a buildfile, update certain targets, load the result into the buildfile, and then resume loading the buildfile.

GNU make has a variant of this functionality: if the makefile specified with the include directive does not exist or is out of date, make will attempt to update it. There are, however, two unfortunate properties of how this works: Firstly, make doesn't stop and update the makefiles when it encounters the include directives. Instead, it ignores non-existing or load outdated included makefiles and continues evaluating until the end, and only then it tries to update them. This means that our makefiles need to be prepared to handle the case where the probe results are not yet known or are outdated. Secondly, if any of the included makefiles were updated, make restarts the process of loading the makefiles from scratch. This can impose a substantial performance penalty on larger projects.

In build2 we've implemented "proper" support for update during load without any of these drawbacks. Specifically, we stop evaluating the buildfile, update all the relevant targets, load them, and continue loading without any restarts. We used this functionality to implement configuration probing as part of the main build with satisfying results (see below for some performance numbers).

Once you get update during load support in your build system, you tend to start uncovering various needs to discover and communicate information back to the build. For example, this functionality can be used to extract the C or C++ compiler predefined macros (predefs) and make them available as variables when evaluating buildfiles.

Before we try to tackle the brittleness issue, let's discuss another relevant detail. The way autoconf and CMake implement function probes is by compiling and linking a test program. They also don't rely on the presence of the function declaration in any header, rather declaring it themselves. In other words, what they really check for is the presence of the corresponding symbol in a library. This approach has a long list of corner cases and drawbacks: The function might be inline or a compiler builtin (and thus without a symbol). The symbol may be present but the function declaration might not be enabled in the corresponding header. Or the function signature might not match what we expect, rendering our call sites invalid.

To give a concrete example, from glibc 2.38 a probe with its own strlcpy() declaration links fine even if compiled without _GNU_SOURCE. But the strlcpy() declaration in <string.h> is only enabled if this macro is defined during compilation.

An alternative approach to checking for the presence of a library symbol would be to obtain the declaration by including the standard header and check whether the call site compiles. This approach doesn't have any of the corner cases listed above. It also closely matches how the function will be used in the actual code. It does require disabling (deprecated) implicit function declarations when compiling C probes, but that's not difficult to do for modern C compilers. As a result, my recommendation is to use the call site compilation for function probes. One additional advantage of this approach is that we can use -fsyntax-only with GCC and Clang (/Zs for MSVC) to speed things up substantially.

Solving the brittleness problem is challenging. In a nutshell, we need to distinguish the failure caused by the absence of the feature we are probing from all other failures. Doing it directly would require analyzing compiler diagnostics, which, I hope, you can see as clearly hopeless.

The problem with analyzing diagnostics is that there are many different compilers and they may change the diagnostics wording even between versions. There are also many ways a probe may fail that would indicate the absence of a feature: header is missing, function declaration is missing, parameter/argument mismatch (in all kind of ways), return value mismatch, etc. So you are looking at maintaining a list of diagnostics patterns for an ever growing list of compilers/versions.

One way to improve your prospects would be to somehow limit yourself only to one compiler version. Maybe this is what the recently open-sourced EDG compiler frontend could be useful for?

The next best thing we can try is to have a "control" probe. The idea is to write a variant of the original probe that we expect to fail in all the same circumstances except when the feature we are interested in is absent. This control should mimic the original as close as possible: it should include the same headers, use the same language constructs, have the same logic, etc. In fact, it is best to have both variants implemented in the same source file. Here is what a probe for strlcpy() could look like:

#include <string.h>

size_t f (void)
{
  char dst[8];

#ifndef CONTROL
  size_t n = sizeof (dst);
  size_t r = strlcpy (dst, "strlcpy", n);
#else
  strcpy (dst, "strlcpy");
  size_t r = 7;
#endif

  return r;
}

Putting it all together, running the probe would then involve two steps:

  1. Compile the control probe passing through any diagnostics and failing if the compilation fails. At the same time extract the header dependency information.
  2. Compile the actual probe ignoring any diagnostics. If the compilation succeeds, assume the feature is present, otherwise – absent.

While the control idea might seem like a clever solution, it's not without drawbacks. The main one is that writing a good control might be challenging. For strlcpy() it is pretty easy to implement a very close control using strcpy(). This makes sure the headers we include are present and usable, the language syntax and logic we use are correct, etc. In other situations writing a close control might be more difficult.

We have tested all these improvements with build2 and you can find the HOWTO article and an example that goes into more detail. We have also evaluated the performance of the overall approach: the overhead of running 500 probes (including controls) on modern hardware (such as Intel i9-12900K) is about half a second.