# Create separate make command for each of CMAKE\_CUDA\_ARCHITECTURES?

**URL:** https://discourse.cmake.org/t/create-separate-make-command-for-each-of-cmake-cuda-architectures/3777
**Category:** Development
**Tags:** lang:cuda
**Created:** [July 18, 2021, 11:42am UTC](https://discourse.cmake.org/t/create-separate-make-command-for-each-of-cmake-cuda-architectures/3777 "2021-07-18T11:42:18Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![cgorac](https://discourse.cmake.org/letter_avatar_proxy/v4/letter/c/c57346/32.png) [@cgorac](https://discourse.cmake.org/u/cgorac)
#### Post date: [July 18, 2021, 11:42am UTC](https://discourse.cmake.org/t/create-separate-make-command-for-each-of-cmake-cuda-architectures/3777/1 "2021-07-18T11:42:18Z")

</div>

Hi,

The CMAKE\_CUDA\_ARCHITECTURES at the moment works so that single target is generated into the make file, where CUDA compiler invocation would list all the architectures desired. So for example, if I put following in my CMake file:

```auto
set_target_properties(tgt PROPERTIES CUDA_ARCHITECTURES "70;75;80")

```

then the build command for CUDA files for this target will look like:

```auto
nvcc ... --generate-code=arch=compute_70,code=[compute_70,sm70] --generate-code=arch=compute_75,code=[compute_75,sm_75] --generate-code=arch=compute_80,code=[compute_80,sm_80] ...

```

The CUDA compiler then generates code serially for each given architecture.

We have a project with couple large CUDA files, that are main culprit for slow build times. From CUDA 11.x, the CUDA compiler has a flag to use multiple threads, for building for different architectures in parallel. However, it’s hard to match value for “-j” make option with the value of this flag in order to avoid oversubscribing CPU cores on build machine. So I’m wondering, would it be possible for CMake to have a flag that would make it possible to change the default behavior and, for above example, generate three targets instead, one per each of the requested architectures? This way, building for different architectures would go in parallel automatically.

Thanks.

---

<div class="post-metadata">

### Author: ![ben.boeckel](https://discourse.cmake.org/letter_avatar_proxy/v4/letter/b/ea5d25/32.png) [@ben.boeckel](https://discourse.cmake.org/u/ben.boeckel)
#### Post date: [July 18, 2021, 12:23pm UTC](https://discourse.cmake.org/t/create-separate-make-command-for-each-of-cmake-cuda-architectures/3777/2 "2021-07-18T12:23:17Z")

</div>

I don’t think CMake’s compilation model would support this well.

As for the `-j` issue, the `Ninja` generator has the concept of “pools” to limit parallelism for specific rules. You could use this to limit CUDA compilations to a pool of size `max(1, nproc / ncuda_arches)` at least. There is no such mechanism in Makefiles though beyond the `-l` flag to have `make` notice a busy machine and avoid launching new processes in such a scenario.

@robert.maynard
