Skip to content

Shorten int and float allocation and deallocation - #8799

Draft
moreal wants to merge 1 commit into
RustPython:mainfrom
moreal:range-int-alloc
Draft

moreal wants to merge 1 commit into
RustPython:mainfrom
moreal:range-int-alloc

Conversation

@moreal

@moreal moreal commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

One of checkbox below must be checked.

  • I did not use AI to write the code of this patch.
  • This PR follows our AI policy

Summary

Every for i in range(...) iteration (above the small-int cache) and every float result creates a fresh object and later frees it. For exact int/float instances with nothing attached, both steps went through the generic object paths. This PR adds short paths for them:

  • Freelist in place: the per-thread freelists were a Cell<FreeList> moved out and back on each push/pop, copying the Vec twice. FreeList::push_local/pop_local edit it in place. Nothing inside can re-enter. int and float use them.
  • Exact-type constructor: Context::new_int/new_float use the new PyRef::new_exact_ref. It skips the dict/heap-type checks, the type-reference clone and the type swap that new_ref needs for an arbitrary class. default_dealloc only pushes exact base-class instances, so a reused husk already has the right type (debug-asserted).
  • Freelist dealloc: payloads with a freelist and no tp_clear (int, float, complex, range) get freelist_dealloc in their vtable. It checks for an exact, untracked, unpublished instance with no __del__, weakrefs, dict or member slots and pushes it straight to the freelist. Anything else falls through to default_dealloc, which in that case would only have done the same push.
  • PyFloat::into_pyobject now goes through new_float, as PyInt already does through new_int.

Why it is faster: the object lifecycle for these types does fewer instructions: no Vec moves, fewer type/flag checks, and none of default_dealloc's __del__/weakref/trashcan/clear machinery or its heavier prologue. callgrind (Linux arm64, per loop iteration):

loop body before after
for i in range(n): pass 470 407 (-13%)
a = x * 1.5 1040 928 (-11%)

pyperformance (macOS arm64, --fast, vs. upstream 43a595ed6): scimark 0.91x, nbody 0.92x, spectral_norm 0.92x, crypto_pyaes 0.95x, float/raytrace/meteor_contest 0.97x. The median across 26 benchmarks is 1.00x. Nothing was more than 2% slower when re-measured interleaved. json_dumps and deepcopy came out at +1.9%, within noise.

extra_tests/snippets/vm_freelist.py pins the behavior that must not change: recycled values and types, __del__ and weakrefs on subclasses (slow path), and subclass husks never coming back as plain ints. It passes before and after this change and on CPython 3.14.

🤖 Generated with Claude Code

Creating and destroying int and float objects went through the generic
object paths even for exact instances with nothing attached:

- The per-thread freelist lived in a `Cell` that was moved out and back
  on every push and pop. Edit it in place (`FreeList::push_local` /
  `pop_local`); nothing inside can re-enter.
- `Context::new_int` / `new_float` now use `PyRef::new_exact_ref`, which
  skips the dict, heap-type and type-swap checks `new_ref` needs for an
  arbitrary class: `default_dealloc` only pushes exact base-class
  instances, so a reused husk already has the right type.
- Payloads with a freelist and no `tp_clear` (int, float, complex,
  range) get `freelist_dealloc`, which pushes an exact, untracked,
  unpublished instance with no `__del__`, weakrefs, dict or member slots
  straight onto the freelist and hands anything else to `default_dealloc`.
- `PyFloat::into_pyobject` goes through `new_float`, as `PyInt` already
  does through `new_int`.

callgrind (arm64) per iteration: empty `for i in range` loop 470 -> 407
instructions, `x * 1.5` loop 1040 -> 928. pyperformance: scimark 0.91x,
nbody 0.92x, spectral_norm 0.92x, crypto_pyaes 0.95x.

Assisted-by: Claude Code:claude-opus-5-5
@coderabbitai

coderabbitai Bot commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the z-ca-2026 Tag to track Contribution Academy 2026 label Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run:pyperformance z-ca-2026 Tag to track Contribution Academy 2026

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant