GPU: materialize flat constant arrays on device - #9434
Conversation
Merging this PR will degrade performance by 11.78%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | words_gather_scalar[65536] |
8.2 µs | 9.4 µs | -12.45% |
| ❌ | WallTime | words_gather_dispatch[1024] |
8 ns | 9 ns | -11.11% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing joe/gpu-flat-constant-arrays (ee5e383) with develop (b825c4f)
Footnotes
-
46 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
171fc0f to
32e03e6
Compare
e271be1 to
1338a6a
Compare
1338a6a to
559fa98
Compare
559fa98 to
4bbfa3c
Compare
Signed-off-by: Joe Isaacs <2413449+joseph-isaacs@users.noreply.github.com>
4bbfa3c to
ee5e383
Compare
Independent of the other follow-up PRs; based directly on #9147. This draft contains one reviewable GPU correctness or performance concern.
What
Why this is needed
Constant nodes are normal compressor output: constant columns, all-null columns, and constant patch values all produce
ConstantArray. The base CUDA executor only handles non-null numeric/decimal constants, so Boolean, UTF-8, binary, extension, and null constants fall back to CPU inside an otherwise GPU-resident tree. The profile exposed a concrete “Nullable UTF-8 Constant CPU fallback” costing 14.573 ms, making this both a correctness-coverage and measured performance fix.Validation
Checked independently against this PR's current base:
cargo check -p vortex-cuda --all-featuresThe combined implementation also passed Rust/CUDA formatting, 164 focused dynamic-dispatch tests, and 28 focused constant-array tests.