Skip to content

Promote Qwen3 audio encoder, S3Gen, PReLU, ESPnet subsampling, and probability segmentation to the framework - #855

Merged
0xShug0 merged 12 commits into
mainfrom
feature/framework-refactor
Oct 10, 2026
Merged

0xShug0 merged 12 commits into
mainfrom
feature/framework-refactor

Conversation

@0xShug0

@0xShug0 0xShug0 commented Oct 9, 2026

Copy link
Copy Markdown
Owner

Summary

  • Promote configurable Qwen3 audio encoder and S3Gen runtimes with explicit weight bindings and session-owned graph caches.
  • Share PReLU, ESPnet three-stage Conv2D subsampling, and probability segmentation across their model callers.
  • Reuse the existing Transformer encoder block in AST/CED and extend explicit SDPA for unequal value widths and tensor score scaling.
  • Remove redundant model-local implementations and the stale Chatterbox Turbo build dependency. No ggml/kernel, server, CLI, UI, or spec changes.

Validation

  • Debug CLI/server builds passed, including ten custom/core composite checks. Turbo also builds independently without Chatterbox.
  • Interface-cleanup A/B: exact outputs across 50 CUDA longform requests, plus CPU/Vulkan repeated-request coverage and CUDA cloning/voice-conversion lifecycle checks.
  • No persistent end-to-end slowdown established in those comparisons. Metal/HIP were not tested.
  • Final CUDA interface-cleanup A/B table: report. Its baseline is identified explicitly; it is not a whole-branch-versus-main benchmark.

@0xShug0
0xShug0 merged commit 2209425 into main Oct 10, 2026
6 checks passed
@0xShug0
0xShug0 deleted the feature/framework-refactor branch October 10, 2026 00:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant