Skip to content

Per packet alloc - #1257

Open
kierank wants to merge 7 commits into
Upipe:masterfrom
kierank:per-packet-alloc
Open

kierank wants to merge 7 commits into
Upipe:masterfrom
kierank:per-packet-alloc

Conversation

@kierank

@kierank kierank commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

No description provided.

The duplicate-packet check only needs the bytes of the previous packet.
The uref_dup and its free cost four pool allocations and releases per
TS packet; a copy of at most 184 bytes costs one extract.
A udict up to 128 bytes now lives inside the pooled udict_inline
structure instead of a separate umem allocation, so allocating, duplicating
and freeing the common small dictionary costs one pool operation instead
of two.  Larger dictionaries, and growth past the embedded space, use the
umem allocator as before.  The test duplicates a small dictionary and grows
both copies past the embedded space.
Each TS packet of a datagram cost a map and an unmap; one map of the
whole datagram checks them all, and the packet by packet path stays for
non-contiguous input or a bad sync word.  Also frees a trailing partial
packet, which was leaked.
…ting it

The extracted uref was a uref_dup of the whole stream, one ubuf per
segment, truncated afterwards and the originals freed on consume.  From a TS
input an access unit has one segment per TS packet, so that was three
pool operations per segment per access unit.  The extracted uref now
takes over the segments up to the boundary, only the boundary segment is
duplicated, and the remainder is attached to the next uref.
Large buffers, such as uncompressed video frames of several MiB, are
typically written and read in full once per frame; on 4 KiB pages every
pass takes a TLB miss per page, and each fresh buffer faults in a page at
a time.  Allocate buffers of 2 MiB and above with 2 MiB alignment and
madvise(MADV_HUGEPAGE) so the kernel backs them with huge pages where
THP is set to madvise or always.  Smaller classes, the free path and the
pool logic are unchanged; the memory still comes from malloc and is
released with free().
A flow definition never has a block, and every input function asks
every data uref whether it is a flow definition, which is a full scan of
its dictionary for an attribute a data uref never has.  The getter now
answers from the ubuf pointer first.  The attribute's accessors are
written out in place of the generated set; set, delete, copy, match and
cmp are unchanged.
…le time

The interleaved copy called memcpy for every sample of every input with
a runtime size, 8 bytes for 32 bit stereo, which is most of the
function's time in a profile.  The common block sizes (4, 8 and 16
bytes) are now copies of a constant size, which the compiler turns into
one or two moves; other sizes keep the plain memcpy.  The bytes written
are unchanged; the existing interleaved test covers the path.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant