Conversation
Creating and destroying int and float objects went through the generic object paths even for exact instances with nothing attached: - The per-thread freelist lived in a `Cell` that was moved out and back on every push and pop. Edit it in place (`FreeList::push_local` / `pop_local`); nothing inside can re-enter. - `Context::new_int` / `new_float` now use `PyRef::new_exact_ref`, which skips the dict, heap-type and type-swap checks `new_ref` needs for an arbitrary class: `default_dealloc` only pushes exact base-class instances, so a reused husk already has the right type. - Payloads with a freelist and no `tp_clear` (int, float, complex, range) get `freelist_dealloc`, which pushes an exact, untracked, unpublished instance with no `__del__`, weakrefs, dict or member slots straight onto the freelist and hands anything else to `default_dealloc`. - `PyFloat::into_pyobject` goes through `new_float`, as `PyInt` already does through `new_int`. callgrind (arm64) per iteration: empty `for i in range` loop 470 -> 407 instructions, `x * 1.5` loop 1040 -> 928. pyperformance: scimark 0.91x, nbody 0.92x, spectral_norm 0.92x, crypto_pyaes 0.95x. Assisted-by: Claude Code:claude-opus-5-5
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One of checkbox below must be checked.
Summary
Every
for i in range(...)iteration (above the small-int cache) and every float result creates a fresh object and later frees it. For exactint/floatinstances with nothing attached, both steps went through the generic object paths. This PR adds short paths for them:Cell<FreeList>moved out and back on each push/pop, copying theVectwice.FreeList::push_local/pop_localedit it in place. Nothing inside can re-enter. int and float use them.Context::new_int/new_floatuse the newPyRef::new_exact_ref. It skips the dict/heap-type checks, the type-reference clone and the type swap thatnew_refneeds for an arbitrary class.default_dealloconly pushes exact base-class instances, so a reused husk already has the right type (debug-asserted).tp_clear(int, float, complex, range) getfreelist_deallocin their vtable. It checks for an exact, untracked, unpublished instance with no__del__, weakrefs, dict or member slots and pushes it straight to the freelist. Anything else falls through todefault_dealloc, which in that case would only have done the same push.PyFloat::into_pyobjectnow goes throughnew_float, asPyIntalready does throughnew_int.Why it is faster: the object lifecycle for these types does fewer instructions: no
Vecmoves, fewer type/flag checks, and none ofdefault_dealloc's__del__/weakref/trashcan/clear machinery or its heavier prologue. callgrind (Linux arm64, per loop iteration):for i in range(n): passa = x * 1.5pyperformance (macOS arm64,
--fast, vs. upstream43a595ed6): scimark 0.91x, nbody 0.92x, spectral_norm 0.92x, crypto_pyaes 0.95x, float/raytrace/meteor_contest 0.97x. The median across 26 benchmarks is 1.00x. Nothing was more than 2% slower when re-measured interleaved. json_dumps and deepcopy came out at +1.9%, within noise.extra_tests/snippets/vm_freelist.pypins the behavior that must not change: recycled values and types,__del__and weakrefs on subclasses (slow path), and subclass husks never coming back as plain ints. It passes before and after this change and on CPython 3.14.🤖 Generated with Claude Code