Skip to content

Speed up SpriteList swap, insert, setitem, and buffer uploads - #2908

Merged
pvcraven merged 1 commit into
developmentfrom
perf/spritelist-swap-insert-upload
Oct 1, 2026
Merged

pvcraven merged 1 commit into
developmentfrom
perf/spritelist-swap-insert-upload

Conversation

@pvcraven

@pvcraven pvcraven commented Oct 1, 2026

Copy link
Copy Markdown
Member

Summary

Speeds up SpriteList.swap(), insert(), __setitem__, and GPU buffer uploads. These were found while reviewing the sprite drawing code.

swap(): O(n) → O(1)

swap() found each sprite's place in the draw order with _sprite_index_data.index(slot), a linear search, even though it's given the positions. The index buffer is always in the same order as the sprite list, so it now swaps index_data[index_1] and index_data[index_2] directly. Negative indexes are made positive first, because the index buffer is longer than the list.

insert() and __setitem__: dictionary membership instead of a list scan

insert() checked sprite in self.sprite_list, and __setitem__ called self.sprite_list.index(sprite). Both now check the sprite_slot dict, the same way append() does.

__setitem__ also used to compare the existing position with the given index, so sprite_list[-1] = sprite_list[-1] raised "already in the list". It now checks whether the sprite is the one being replaced, which also fixes that.

Upload only the slots in use

Any change to a sprite re-uploaded the whole capacity of the affected buffer, so cost grew with capacity rather than sprite count. Capacity doubles as lists grow and never shrinks after pop() or remove(). write_sprite_buffers_to_gpu() gets two optional arguments, slot_count and index_count. The buffer backend uses them to write zero-copy memoryview slices of just the slots in use.

The rest of each orphaned buffer stays undefined, which is safe: the index buffer only refers to slots below slot_count. The texture (WebGL) backend accepts the arguments but still writes whole textures. Both arguments default to None (write everything), so existing callers of this method don't change.

Benchmarks

Old and new code alternated, 3 rounds, best result with the range in brackets:

Case Before After
swap(-1, -2) on a 10,000 sprite list 222 µs (222–227) 0.22 µs
sl[9000] = sprite on a 10,000 sprite list 64.9 µs (65–75) 1.46 µs
insert(5000) + pop(5000) on a 10,000 sprite list 77.3 µs (77–93) 22.9 µs
100 sprites, capacity 256: move one, draw 89.2 µs (89–106) 57.3 µs
100 sprites, capacity 65,536: move one, draw 213 µs (213–236) 94.2 µs
5,000 sprites: move all, draw 3723 µs (3723–3807) 3061 µs (3061–3741)

The last row is the normal "everything moves" case. The ranges overlap, so I'd call it unchanged within noise, but it's not slower. insert() is still O(n) because inserting into a Python list and array is, but the extra scan is gone.

Tests

  • test_swap_draw_order: swaps including negative, mixed, and same-index cases. It checks the list and the draw order read back from the GPU index buffer.
  • test_swap_out_of_range: IndexError, and nothing changes.
  • test_setitem_negative_index: setting a sprite to its own position, adding a sprite that's already elsewhere (raises), an index out of range (raises), and replacing at a negative index.
  • test_insert_already_in_list.
  • test_gpu_buffers_match_after_changes: 200 random pops, appends that reuse freed slots, inserts, swaps, and moves, with buffer growth, drawing every 10 steps. At the end, the draw order and each sprite's position, depth, angle, and size are compared against what was uploaded to the GPU. With the position upload deliberately cut one slot short, this test fails.

On development, only test_setitem_negative_index fails, since that's the one behavior change. Full suite on pyglet 3.0.dev11: 1419 passed. The 3 failures are the render tests that only fail on my machine. Ruff is clean, and mypy reports no errors in sprite_list.py.

🤖 Generated with Claude Code

- swap(): the index buffer is in the same order as the sprite list, so
  use the given positions instead of searching it with .index(). O(1)
  instead of O(n).
- insert() and __setitem__: check membership with the sprite_slot dict
  instead of scanning the list. __setitem__ now also accepts setting a
  negative index to the sprite already there.
- write_sprite_buffers_to_gpu(): new optional slot_count and index_count
  arguments. The buffer backend writes only the slots in use, through
  zero-copy memoryview slices, instead of the whole capacity. The
  texture (WebGL) backend still writes everything.

Measured: swap at the end of a 10k list 222 -> 0.22 us, setitem 65 ->
1.5 us, insert+pop 77 -> 23 us, move one sprite and draw 1.6-2.3x
faster.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@pvcraven
pvcraven merged commit d771686 into development Oct 1, 2026
7 checks passed
@pvcraven
pvcraven deleted the perf/spritelist-swap-insert-upload branch October 1, 2026 20:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant