How to Fix autoconf-style Config Probing
Posted on 8 Oct 2026 by
Boris Kolpackov
Configuration probing as implemented in autoconf/CMake/etc
involves compiling and linking a test program to determine whether a
particular feature, such as a function, is available on the platform being
targeted. For example, we may prefer to use the strl*() family
of functions in our codebase. However, these functions are not (yet)
standard and are not provided by all libc implementations. As a result, we
may wish to detect whether they are present and if not, provide fallback
implementations or use alternatives. One way to do this detection would be
to compile and link a test program that tries to use the functions we are
interested in. If that succeeds, then we conclude the functions are
available.
On the face of it, this approach is appealing. In particular, it is
adaptable in the sense that we don't have to do anything to support
platforms that may not even exist yet. For example, if someone decides to
write yet another libc for Linux, we don't have to do anything to support it
– the existing strl*() probes will sort it out. In fact,
even already released versions of our project will automagically support
this new libc.
This approach does have a few annoying problems. Here are the main ones:
- It is wasteful: There is no need to keep compiling the
strl*()probes on, say, FreeBSD, where these functions were available for eons. At the limit this becomes absurd, like keep probing for a feature while the latest target that doesn't have it would not even be able to perform the probe. For a good example, see A Generation Lost in the Bazaar. - It is brittle: We decide that a feature is absent based on the
failure to compile/link a test program. But a lot of other things can lead
to a failure to compile or link: mistakes in the test, misconfigured build,
missing feature test macros such as
_GNU_SOURCE, etc.For example, a lot of weeping and gnashing of teeth was recently caused by false negatives due to sloppily written probes. They stopped compiling because GCC and Clang stopped accepting certain long-deprecated C constructs.
The failure mode is also insidious: a false negative silently leads to the feature not being used, leading to missing functionality, suboptimal performance, etc.
- It is slow: While compiling a single probe doesn't take long,
compiling several hundreds is noticeable. To exacerbate the problem, both
autoconfand CMake do it serially. - It lacks change-tracking: Existing tools (
autoconf, CMake) do not re-run the relevant probes when their inputs change. For example,strl*()were added in glibc 2.38. If we upgraded from 2.37, we would want all the already configured projects on our machine to detect the change and start using the newly available functions.
Solving the first problem (wastefulness) requires a completely different
approach. One alternative is to use what we can call "expectation-based
configuration": we assume a feature is available if certain conditions are
met. For example, for strl*() we could assume these functions
are available if we are targeting FreeBSD or glibc version 2.38 or later (of
course, a complete
implementation would also need to check for other platforms and/or libc
implementations). This approach has been successfully used in build2 on configuration-heavy
projects such as Qt and FFmpeg (see libbuild2-autoconf
for details).
Ok, let's say we still wish to do configuration probing for some reason or for some special cases. Can we solve, or at least mitigate, the remaining problems? Let's save the brittleness problem for last and take a stab at the remaining two: slowness and lack of change-tracking.
A high-level view of what we are doing during configuration probing can be summed up like this: we are compiling and linking a number of test programs, except that the result we are after is not the programs but rather the status: whether the compilation and linking succeeded or failed. We would like to do this in parallel and also keep track of changes to inputs: test source itself, recursive set of headers included by it, compile/link options, etc.
Doesn't the shape of this problem look familiar? What existing problem
requires us to compile and link a bunch of source files in parallel and with
proper change-tracking? That's right, this is how we build our software with
existing build systems. Even make can do this reasonably
well.
Apparently, CMake generates an individual project per each probe and then runs the underlying build system to build it. But it neither uses this to run multiple probes in parallel nor to track changes.
So couldn't we just use the build system to do the probing? And while at it couldn't we get rid of the whole separate configuration/project generation step?
It could work like this: we run the build system to update our project,
it builds (or re-builds) the probes as necessary and then uses the resulting
information to build our project source code. Specifically to our
strl*() example, the build system would compile and link
strlcpy.c and strlcat.c fallback implementation if
the probes for these functions returned negative results.
Another advantage of using a proper build system for probes is the
ability to establish dependencies between probe results. For example, there
is no use wasting time probing for strl*() if there is no
<string.h>.
There is one snag, though: the results of the probes need to be known
when loading and evaluating the buildfiles. For example, we decide whether
to include strlcpy.c and strlcat.c into the build
while evaluating buildfile definitions. Here is a GNU
make-based illustration:
hello: hello.o hello.o: hello.c ifndef have_strlcpy hello: strlcpy.o strlcpy.o: strlcpy.c else CPPFLAGS += -DHAVE_STRLCPY endif ifndef have_strlcat hello: strlcat.o strlcat.o: strlcat.c else CPPFLAGS += -DHAVE_STRLCAT endif
In the above example, have_strlcpy and
have_strlcpy would need to be known when make is
evaluating this makefile but if running the corresponding probes is part of
the overall build, then their values are only known later, once the makefile
has been evaluated and make starts actually building the
targets.
We could easily overcome this snag if we had the ability to pause loading a buildfile, update certain targets, load the result into the buildfile, and then resume loading the buildfile.
GNU make has a variant of this functionality: if the makefile specified
with the include directive does not exist or is out of date,
make will attempt to update it. There are, however, two
unfortunate properties of how this works: Firstly, make doesn't
stop and update the makefiles when it encounters the include
directives. Instead, it ignores non-existing or load outdated included
makefiles and continues evaluating until the end, and only then it tries to
update them. This means that our makefiles need to be prepared to handle the
case where the probe results are not yet known or are outdated. Secondly, if
any of the included makefiles were updated, make restarts the
process of loading the makefiles from scratch. This can impose a substantial
performance penalty on larger projects.
In build2 we've implemented "proper" support for update
during load without any of these drawbacks. Specifically, we stop
evaluating the buildfile, update all the relevant targets, load them, and
continue loading without any restarts. We used this functionality to
implement configuration probing as part of the main build with satisfying
results (see below for some performance numbers).
Once you get update during load support in your build system, you tend to start uncovering various needs to discover and communicate information back to the build. For example, this functionality can be used to extract the C or C++ compiler predefined macros (predefs) and make them available as variables when evaluating buildfiles.
Before we try to tackle the brittleness issue, let's discuss another
relevant detail. The way autoconf and CMake implement function
probes is by compiling and linking a test program. They also don't rely on
the presence of the function declaration in any header, rather declaring it
themselves. In other words, what they really check for is the presence of
the corresponding symbol in a library. This approach has a long list of
corner cases and drawbacks: The function might be inline or a compiler
builtin (and thus without a symbol). The symbol may be present but the
function declaration might not be enabled in the corresponding header. Or
the function signature might not match what we expect, rendering our call
sites invalid.
To give a concrete example, from glibc 2.38 a probe with its own
strlcpy() declaration links fine even if compiled without
_GNU_SOURCE. But the strlcpy() declaration in
<string.h> is only enabled if this macro is defined during
compilation.
An alternative approach to checking for the presence of a library symbol
would be to obtain the declaration by including the standard header and
check whether the call site compiles. This approach doesn't have any of the
corner cases listed above. It also closely matches how the function will be
used in the actual code. It does require disabling (deprecated) implicit
function declarations when compiling C probes, but that's not difficult to
do for modern C compilers. As a result, my recommendation is to use the call
site compilation for function probes. One additional advantage of this
approach is that we can use -fsyntax-only with GCC and Clang
(/Zs for MSVC) to speed things up substantially.
Solving the brittleness problem is challenging. In a nutshell, we need to distinguish the failure caused by the absence of the feature we are probing from all other failures. Doing it directly would require analyzing compiler diagnostics, which, I hope, you can see as clearly hopeless.
The problem with analyzing diagnostics is that there are many different compilers and they may change the diagnostics wording even between versions. There are also many ways a probe may fail that would indicate the absence of a feature: header is missing, function declaration is missing, parameter/argument mismatch (in all kind of ways), return value mismatch, etc. So you are looking at maintaining a list of diagnostics patterns for an ever growing list of compilers/versions.
One way to improve your prospects would be to somehow limit yourself only to one compiler version. Maybe this is what the recently open-sourced EDG compiler frontend could be useful for?
The next best thing we can try is to have a "control" probe. The idea is
to write a variant of the original probe that we expect to fail in all the
same circumstances except when the feature we are interested in is absent.
This control should mimic the original as close as possible: it should
include the same headers, use the same language constructs, have the same
logic, etc. In fact, it is best to have both variants implemented in the
same source file. Here is what a probe for strlcpy() could look
like:
#include <string.h>
size_t f (void)
{
char dst[8];
#ifndef CONTROL
size_t n = sizeof (dst);
size_t r = strlcpy (dst, "strlcpy", n);
#else
strcpy (dst, "strlcpy");
size_t r = 7;
#endif
return r;
}
Putting it all together, running the probe would then involve two steps:
- Compile the control probe passing through any diagnostics and failing if the compilation fails. At the same time extract the header dependency information.
- Compile the actual probe ignoring any diagnostics. If the compilation succeeds, assume the feature is present, otherwise – absent.
While the control idea might seem like a clever solution, it's not
without drawbacks. The main one is that writing a good control might be
challenging. For strlcpy() it is pretty easy to implement a
very close control using strcpy(). This makes sure the headers
we include are present and usable, the language syntax and logic we use are
correct, etc. In other situations writing a close control might be more
difficult.
We have tested all these improvements with build2 and you
can find the HOWTO
article and an example
that goes into more detail. We have also evaluated the performance of the
overall approach: the overhead of running 500 probes (including controls) on
modern hardware (such as Intel i9-12900K) is about half a second.