make BMI feature request

The idea would be similar to make codegen. This allowed a minimal set of actions to be carried out from the build process, so that static checkers can nicely do their job, this is way better compared to the alternative of build all first.

Today a new similar problem is present, when using modules, clang-tidy does support modules, but requires the BMIs to be available. Closest to getting this done today, I think, is again build all.

Would it be possible to have a dedicated target “BMI” that creates all the BMIs in the current project, so that static analyzers can pick up after that, this should be again way smaller compared to build all … assuming that the static analyzers can nicely pick them up after this stage.

What do you think ?

Cc: @ben.boeckel @vito.gamberini

We could also consider generating BMIs as part of codegen since its purpose is to prepare for static analyzers.

For reference, one can integrate clang-tidy into the build itself using CMAKE_<LANG>_CLANG_TIDY. BMIs should then be available as needed. Of course this also requires fully building all the targets of interest.

We want to support pure BMI generation at some point but it requires a lot of rework internally for how CMake models BMI-producing sources.

Right now the answer is you produce a full build, then all the BMIs will be available.

The reason is CMake strictly segments the world into object-file producing commands and non-object-file producing (BMI-only) commands (called “synthetic”). There is no way today to translate between these.

If the BMI you need comes from a command which also produces an object file (which is the common case), there is no way to “extract” only the BMI-parts and run them. We must produce the object file too, and that’s just a full build.

fair point, though I can see reasons why it still may be much smaller than a full build:

  • not all libraries are modules yet
  • a library modeled as a module might have several implementation units, meaning “PMIU → BMI + object” is only a part of it, off course the same applies to creating the internal BMI for module partitions

Agreed, and I can pull out “the BMI producing commands of this target”, but then I need to run the collation step to run them (because they also produce object files, which require correct ordering to be built).

The collation step lives in the bowels of the pipeline and is not easily disentangled. It would be non-trivial work to do so. If I had BMI-only command equivalents of the object-producing commands I wouldn’t need to worry about collation at all.

So yes we should do it, but the “shortcut” solution is actually quite involved, and suboptimal. And the complete solution is equally tricky.

Ben talked about this early on. It was known CMake’s architecture for supporting modules would be pretty bad for intellisense and static analysis, and effectively require a full build. It was argued that shortest path to supporting modules at all was worth it, and afterwards we could second-system the choices to better support auxiliary tooling.

So yes, we should do it, we will do it. It’s on the menu, it’s just not a 500 line patch I can drop in tomorrow.

(Blame Fortran, which C++20 module support is based on, but doesn’t have BMIs or extensive auxiliary tooling)

thanks for the insights, well it gives us something to look forward too :blush:

anyway many thanks for all the continuous improvements to cmake.

One step at a time, we come a long way for modules, so not all can happen at once.

Something I see at almost every consulting client with a non-trivial-sized project is the clang-tidy job eventually ends up being the longest running job in CI. And that’s without any co-compilation, that’s just purely running clang-tidy on its own, usually via the run-clang-tidy script that runs clang-tidy in parallel across the contents of a compile_commands.json. Relying on co-compilation for clang-tidy in CI would be a non-starter in most organisations because of the increase in the long pole job time, which in turn limits PR throughput.

On the plus side, such clang-tidy jobs tend to be embarrassingly parallel. That means you get very good bang-for-your-buck by throwing more CPU at it, as long as the CMake configure time and any code generation time is small (or also embarrassingly parallel). If the project uses C++20 modules, this may complicate that picture, depending on how efficiently your project can produce the BMIs that clang-tidy needs. Having to build the project to get those BMIs means now your build needs to be highly parallel with no bottlenecks if you want to continue to get good return for throwing more CPUs at the problem. That’s not a given, especially once you start talking about 32 or more CPUs being available.

To be frank, static analysers are my biggest argument against projects adopting C++20 modules right now, precisely because of the above problem. There’s a similar situation with other clang-based tools like iwyu. Projects that use such static analysis tools are not a good fit for C++20 modules, at least not until producing BMIs is trivially fast, and I suspect that’s inherently challenging, if not impossible, due to how modules work.