-
Notifications
You must be signed in to change notification settings - Fork 104
Pull requests: InfiniTensor/InfiniLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat(engine): integrate prefix eviction and chunked parallel execution
#573
opened Sep 16, 2026 by
big-hip
Loading…
32 of 41 tasks
feat: add priority-aware request scheduling
#571
opened Sep 13, 2026 by
xiaoba17
Loading…
26 of 36 tasks
fix(moore): enable Mate flash-attn inference
#567
opened Sep 10, 2026 by
spike-zhu
Collaborator
Loading…
1 of 49 tasks
feat(nvidia): add reusable GGUF Route B support for Qwen3.5
#559
opened Sep 3, 2026 by
xindongliu594
Loading…
32 of 48 tasks
perf: fuse unquantized Qwen3-Next GDN projections
#555
opened Sep 2, 2026 by
T4t4KAU
Loading…
26 of 39 tasks
feat(server): add agent support with tool-call and reasoning parsing
#554
opened Sep 1, 2026 by
rubik-hua
Contributor
Loading…
feat: support Ktransformers, CPU-GPU MoE offload via FusedMoE layer
#548
opened Aug 21, 2026 by
whjthu
Contributor
Loading…
37 of 49 tasks
fix(cuda-graph): keep replay metadata dynamic across tensor-parallel ranks
#540
opened Aug 15, 2026 by
junjiewang253-ctrl
Loading…
feat: add aclnnMatmulAllReduce fusion in InfiniLM for Ascend RowParallelLinear
#533
opened Aug 11, 2026 by
ShaneWoof
Contributor
Loading…
feat(hygon): add Qwen3-235B-A3B BF16/W8A8 inference support
#532
opened Aug 7, 2026 by
qinyiqun
Contributor
Loading…
49 tasks
feat(engine): overlap decode steps with asynchronous token handoff
#524
opened Aug 3, 2026 by
qinyiqun
Contributor
Loading…
feat: add Qwen3.6 MoE model support
#521
opened Jul 31, 2026 by
qinyiqun
Contributor
Loading…
49 tasks
perf(server): coalesce streaming SSE output
#517
opened Jul 28, 2026 by
wooway777
Collaborator
Loading…
49 tasks
Previous Next
ProTip!
Follow long discussions with comments:>50.