gh-156250: Speed up inherited mapping assignment slot calls - #156258
Open
koxudaxi wants to merge 2 commits into
Open
gh-156250: Speed up inherited mapping assignment slot calls#156258koxudaxi wants to merge 2 commits into
koxudaxi wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
__setitem__and__delitem__share themp_ass_subscriptslot. If a Python subclass of a C type overrides only one of them, the inherited operation is still dispatched throughslot_mp_ass_subscript()and called through awrapper_descriptor.When lookup returns the expected unbound wrapper descriptor,
slot_mp_ass_subscript()now calls the wrapped C slot directly. Python overrides and other descriptors continue through the generic call path.Benchmark
I found this while profiling Counter-heavy code in a private project.
Counteroverrides__delitem__but inheritsdict.__setitem__, so several Counter operations go through this path.I compared the feature commit
8ca6cd042b1e1fe91e2751c5abf3204631c749c0with its direct parentb062727097e997bcb900e11503d3248daac903da. Both were release builds on macOS 15.7.1 arm64 with the same PGO/LTO configuration.c[key] = valuec[key] += 1c.update(mapping)c.subtract(mapping)a + ba | bCounter(iterable)andCounter.update(iterable)are not affected. They already use the_count_elements()C fast path.I also checked larger workloads with the same two builds. NetworkX
nx.triangles()was 5.1% to 9.0% faster across four graph shapes. The BPE benchmark in pyperformance was 6% to 8% faster in repeated runs.The same one sided
__setitem__/__delitem__pattern also exists outside Counter, but most of the other workloads I checked showed little or no end to end difference.Given these results, I think this optimization is worth the extra complexity in
slot_mp_ass_subscript(). The affected pattern is fairly narrow, but the improvement shows up in normal Counter operations and in larger workloads.(Updated on August 26, 2026 with additional Counter and real world benchmark results.)
Tests
Tests cover asymmetric
dictsubclasses, inherited slice deletion on alistsubclass, and fallback for custom or incompatible descriptors. The targeted tests pass on regular and free-threaded debug builds.Fixes #156250