Skip to content

Fix device placement for memory-planned buffers in Runtime.load_program - #22058

Open
shoumikhin wants to merge 13 commits into
mainfrom
fix-pybindings-planned-buffer-device
Open

Fix device placement for memory-planned buffers in Runtime.load_program#22058
shoumikhin wants to merge 13 commits into
mainfrom
fix-pybindings-planned-buffer-device

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

Fix device placement for memory-planned buffers in Runtime.load_program

Replaces #22057. That pull request was force-pushed to a commit with no
history in common with main, which made GitHub close it permanently. Same
branch, same change, correct history.

The problem

ExecuTorch has two ways to load a model from Python. With the CUDA backend, one
of them works and the other crashes. Same .pte file, same .ptd weights file,
same default export settings. Only the loader differs.

# works
_load_for_executorch(pte_path, ptd_path).forward([x])

# crashes
Runtime.get().load_program(pte_path, data_path=ptd_path) \
    .load_method("forward").execute([x])

The crash looks like this:

[cuda_backend.cpp:548] Tensor 0 has device_type=CUDA but its data pointer
0x... is not backed by CUDA device memory
(cudaPointerGetAttributes err=0, cudaMemoryType=0).
RuntimeError: method->execute() failed with error 0x12

Why it happens

A .pte file records where each memory-planned buffer has to live, on the host
or on an accelerator. That is the non_const_buffer_device field in the plan.

ProgramMemory in extension/pybindings/pybindings.cpp never read that field.
It allocated every planned buffer as host memory (std::vector<uint8_t>). So a
program that asked for device memory received a host pointer. The CUDA backend
checked the pointer, saw it was not device memory, and refused to run.

.pte says:  buffer 0 -> CUDA:0
before:     buffer 0 -> std::vector<uint8_t>        (host)   -> backend rejects
after:      buffer 0 -> DeviceMemoryBuffer::create  (device) -> backend accepts

The fix

extension/module/module.cpp hits the same problem in the C++ Module API and
answers it differently, because it has a share_memory_arenas flag and this
loader has none. Module refuses to load a device-planned method when that flag
is set, and otherwise builds per-method arenas for every method.
Runtime.load_program shares arenas unconditionally today and offers no way to
turn that off, so refusing would stop files loading that load now. It keeps the
shared host arenas for host methods and gives a device-planned method its own.

  1. ProgramMemory now receives the per-buffer device list alongside the sizes,
    and allocates each buffer on the device it is tagged for. Host-tagged buffers
    keep using std::vector<uint8_t>, exactly as before.
  2. Buffer indices are plan local. Buffer 0 of one method has nothing to do with
    buffer 0 of another, so one set of arenas shared by index cannot describe a
    file where one method plans onto the host and another onto an accelerator. A
    method with any device-tagged buffer therefore gets its own arenas, and every
    other method keeps using the shared host arenas.
  3. Those arenas are built in load_method, not when the program is loaded. A
    file containing one accelerator method still opens on a machine without that
    accelerator, still lists its methods, and its host methods still load and
    run. Only loading the accelerator method fails, and it fails naming the
    buffer and the device it could not allocate.
  4. The caller reads the device for each buffer through
    MethodMeta::memory_planned_buffer_device.
  5. Two deprecated calls were replaced with their current names. Both are plain
    forwarders, so behavior is unchanged.

Program.load_method in runtime/__init__.py documents all of this, including
the part that is not new: two host-only methods of the same program share one
set of arenas and therefore overwrite each other's intermediate values.

Existing programs are unaffected

non_const_buffer_device is optional. MethodMeta::memory_planned_buffer_device
returns Device{CPU, 0} when the field is absent, which is the case for
CPU-only programs and for .pte files produced before the field existed. Such a
program keeps the shared arenas, the host allocation path, and the single
argument HierarchicalAllocator, so MemoryManager::has_device_memory() stays
false for it as that constructor documents.

Test plan

Two tests in extension/pybindings/test/test_pybindings.py.

test_program_loads_when_one_method_is_device_planned covers the refusal
path and needs no GPU, so it runs in the existing CPU-only job. It exports a two
method program where one method has a device-tagged planned buffer and the other
has none, then checks that the program loads, that the host method runs, that
loading the device method reaches the device allocator and is refused there, and
that the host method still runs afterwards. It first asserts that the exported
program really does carry a CUDA-tagged planned buffer, so it cannot quietly
degrade into a plain multi-method test if planning stops tagging devices. It
skips itself on a build that links the CUDA backend, because that registers a
CUDA allocator at static init, the registry has no way to drop one, and the
request would then be satisfied with or without this change.

test_device_planned_method_allocates_on_the_device covers what the refusal
is protecting: on a build that does have a device allocator, the arena has to
come off the device rather than out of host memory. It builds a two method
program where one method is lowered to the CUDA backend and the other is not,
then measures free device memory with torch.cuda.mem_get_info around each
load_method call. It asserts that loading the host method takes no device
memory, that loading the device method takes at least 90 percent of the planned
device bytes, that both methods return correct numbers, that running the host
method in between does not disturb the device method, and that releasing the
program returns the memory. It skips unless the build links the CUDA backend and
a device is visible.

That second test can only run where a device allocator is registered, so this
pull request also wires it into the job that has one. .ci/scripts/test-cuda-build.sh
runs it, and .github/workflows/cuda.yml now triggers on changes under
extension/pybindings/ and to runtime/__init__.py. Before this, no job in the
repository built the Python extension with a device allocator and then ran the
pybindings tests, which is exactly why this class of defect was invisible.

Measured

Linux x86_64, NVIDIA A100 80GB, compute capability 8.0, CUDA 13.0, Python 3.12,
torch 2.13.0. Two builds from one source tree, one CPU only and one with the
CUDA backend. The before column is the merge base with main, produced by
swapping only pybindings.cpp and rebuilding, so both columns are the same
machine, the same model files and the same everything else.

Two methods in one program, forward on the host and forward2 lowered to
CUDA, planned device bytes 50331648 (48 MiB):

Measurement Before After
Device memory taken by load_program 0 MiB 0 MiB
Device memory taken by loading the host method 0 MiB 0 MiB
Device memory taken by loading the CUDA method 0 MiB 48 MiB
CUDA method produces correct numbers no, error 0x12 yes
Host method produces correct numbers yes yes
CUDA method still correct after running the host method not reached yes
Device memory returned when the program is released nothing to return 48 MiB

The before column is not a generic failure. The CUDA backend names the defect
itself:

[cuda_backend.cpp:548] Tensor 0 has device_type=CUDA but its data pointer
0x7f5842fff010 is not backed by CUDA device memory
(cudaPointerGetAttributes err=0, cudaMemoryType=0).
[method.cpp:1528] CALL_DELEGATE execute failed at instruction 2: 0x12

Suites, on the same two builds:

Suite CUDA build CPU-only build
extension/pybindings/test/test_pybindings.py 39 passed, 1 skipped, 2 failed 39 passed, 1 skipped, 2 failed
the same file filtered to -k device 4 passed, refusal test skipped 4 passed, success test skipped
test_device_planned_method_allocates_on_the_device alone, run the way the CI script runs it passed skipped

The two failures are test_method_quantized_ops and test_quantized_ops. They
are pre-existing and unrelated: they need the quantized AOT library preloaded,
which the Buck target does and a bare pytest invocation does not. They
reproduce identically at the merge base.

Linux aarch64, Jetson Orin Nano, Python 3.10, CPU-only build: identical counts to
the x86_64 CPU-only column above, including the device tests and the same two
pre-existing quantized-op failures. The CUDA success path is not reachable on
that board, because its GPU needs a PyTorch build pinned to a different version
than this repository requires, so only the host paths and the refusal path are
covered there.

Not covered

  • Leaving a device-planned method out of the shared host arenas is not covered
    by any test. Delete that skip and both tests still pass, because nothing in
    Python can observe the shared arena sizes. What it saves is host memory that
    nothing reads, which grows with the model, so it is worth a C++ test later.
  • The device path is measured on one accelerator, an A100 with compute
    capability 8.0. Not measured on a Jetson board or on any non-CUDA
    accelerator.
  • has_device_buffers and make_method_memory ask MethodMeta for one buffer
    at a time, and MethodMeta::memory_planned_buffer_device scans the sparse
    device list on each call, so the cost is the buffer count times the device
    entry count. Real programs measured here have 2 or 3 buffers and 1 device
    entry, and extension/module/module.cpp already reads the same metadata the
    same way, but both counts come from the file. Removing the concern properly
    means a bulk accessor on MethodMeta, which would fix both callers at once
    and belongs in its own change.

Copilot AI lite review requested due to automatic review settings August 23, 2026 01:08
@pytorch-bot

pytorch-bot Bot commented Aug 23, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/22058

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures, 31 Pending, 1 Unrelated Failure

As of commit 3862c42 with merge base 85ec16e (image):

NEW FAILURES - The following jobs have failed:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 23, 2026
Copilot AI review requested due to automatic review settings August 23, 2026 03:08
@shoumikhin
shoumikhin force-pushed the fix-pybindings-planned-buffer-device branch from a1dd903 to 386752e Compare August 23, 2026 03:08

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

A .pte file records where each memory-planned buffer has to live, on the host
or on an accelerator, in the plan's non_const_buffer_device field. ProgramMemory
never read that field and allocated every planned buffer as host memory, so a
program asking for device memory received a host pointer. With the CUDA backend
this failed at execute time with "not backed by CUDA device memory", while the
same .pte loaded through _load_for_executorch worked.

Allocate each planned buffer on the device it is tagged for. Buffer indices are
plan local, so one set of arenas shared across methods by index cannot describe
a file where one method plans onto the host and another onto an accelerator. A
method with any device-tagged buffer therefore gets its own arenas, and every
other method keeps using the shared host arenas.

extension/module/module.cpp faces the same problem and answers it differently,
because it has a share_memory_arenas flag and this loader has none. Module
refuses to load a device-planned method when that flag is set, and otherwise
builds per-method arenas for every method. Runtime.load_program has no such
flag and shares arenas unconditionally today, so refusing would break loading
files that already load. It keeps the shared host arenas for host methods and
gives a device-planned method its own instead.

Those arenas are built when the method is loaded rather than when the program
is loaded, so a file containing one accelerator method still opens on a machine
without that accelerator, still lists its methods, and its host methods still
load and run.

non_const_buffer_device is optional and MethodMeta reports CPU when it is
absent, so CPU-only and older programs keep the shared arenas, the host
allocation path, and the single argument HierarchicalAllocator that leaves
MemoryManager::has_device_memory() false.

Adds test_program_loads_when_one_method_is_device_planned, which exports a two
method program where one method has a device tagged planned buffer and the other
has none, then checks that the program loads, that the host method runs, that
loading the device method reaches the device allocator and is refused there, and
that the host method still runs afterwards. It needs no GPU, and it skips itself
on a build that links the CUDA backend, since that registers a CUDA allocator
which would satisfy the request. Verified that the test fails when the change to
pybindings.cpp is reverted: the old loader put the device buffer in host memory,
so the method loaded and no error was raised.
@shoumikhin
shoumikhin force-pushed the fix-pybindings-planned-buffer-device branch from 386752e to 58cf1b1 Compare August 23, 2026 04:05
Copilot AI review requested due to automatic review settings August 23, 2026 04:05

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Program.load_method does not treat every method the same way. A method
whose memory planned buffers are all on the host reuses one set of arenas
shared with the other host only methods of the same program, so running a
second such method can overwrite the intermediates and outputs of the
first. A method with a buffer placed on an accelerator gets its own arenas
instead.

Nothing said so. Say it in the docstring users read.
Device allocation is the only step here that can fail, for example when no
allocator is registered for the requested device or when the device is out
of memory. It ran last, so every host arena of that method was already
allocated and zero filled before the failure was reported, and all of it
was then discarded.

Swap the two members so the step that can fail runs first. Member
initialization follows declaration order, so the order is what makes this
work and a reorder would silently undo it. Say that in a comment.
ProgramMemory hands one span per planned buffer and one device tag per
planned buffer to HierarchicalAllocator, which checks that the two counts
agree. That check aborts the process. In a Python extension that means the
interpreter dies with no traceback and nothing the caller can catch.

Both constructors keep the counts equal today, so this is not reachable, but
the invariant lives only in the callers. Check it where the error can still
become a Python exception.
load_method walked every planned buffer once through has_device_buffers to
decide whether the method needs its own arenas, then walked them all again
through make_method_memory to collect the sizes and devices. Each device
lookup scans the program's sparse device list, so the second walk repeats
work the first one already did.

Have make_method_memory return nullptr when every buffer is on the host.
One pass then answers both questions. has_device_buffers stays because the
program constructor still needs only the answer, not the sizes.
pybindings.cpp includes runtime/core/device_memory_buffer.h but the
pybindings targets did not depend on it. The build worked only because the
header came in transitively through extension/module, and the rule opts out
of automatic dependency checking, so nothing would report the omission.
Copilot AI review requested due to automatic review settings August 24, 2026 02:50

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI review requested due to automatic review settings August 24, 2026 05:10

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Copilot AI review requested due to automatic review settings August 24, 2026 05:12

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Three comments describe the code inaccurately.

The first says device buffers are allocated first because they hold the
only allocation that can fail. Allocating a host arena can also fail, so
the real reason is only that a device failure should happen before the
host arenas are allocated and zero filled.

The second implies the buffer and device count guard converts a
reachable abort. Both vectors are filled in lockstep today, so the guard
only fires if a future caller breaks that.

The third says a device planned method inflates arenas nobody will read.
Every host only method reads those arenas. What that method would inflate
is their size.
Result::get() aborts the process if the Result holds an error, so a
corrupt program file that reports a bad planned buffer size kills the
interpreter instead of raising. This line was already rewritten by an
earlier commit in this change, and the device buffer code added next to
it already checks the same call, so the two now behave the same way.
The docstring said a second host only method may overwrite the outputs
of the first. It cannot: outputs are copied out on every call, so what
sharing affects is the intermediate values, not what execute returns.

It also said a failed device allocation affects only that method. There
is no unload, so whatever a device method does claim stays claimed for
every method loaded afterwards. Both points are now stated as they are.
The test file imports DeviceType from executorch.exir.schema and
executorch.exir.backend.test.device_util, and the three test targets
picked both up only through make_test. A direct import needs a direct
dependency, so the targets no longer break if make_test drops either
one.
The existing test covers the case where no device allocator is
registered and the load is refused. Nothing covered the case the change
exists to fix: a build that does have a device allocator, where the
arena has to come off the device.

The new test lowers one of two methods to the CUDA backend, checks that
free device memory drops by the planned buffer size when that method is
loaded, checks that both methods return the right values, and checks
that the memory comes back when the program is released. On the previous
behavior the arena stayed in host memory and the backend rejected the
pointer with an execute error, so this test fails without the fix.

A real accelerator is required, so the test skips unless one is present.
The CUDA build job now runs it, and the CUDA workflow now triggers on
changes under extension/pybindings, which is where this code lives.
Copilot AI review requested due to automatic review settings August 24, 2026 06:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants