Skip to content
View StarWorkshop's full-sized avatar
🏠
Working from home
🏠
Working from home
  • Jinan University
  • Shenzhen, China
  • 01:20 (UTC +07:00)

Block or report StarWorkshop

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
StarWorkshop/README.md

Zheng Han — 语音识别 × 端到端语音大模型

Jinan University speech recognition end-to-end spoken LM streaming-native offline-first

研究旅程:wavrail → dialogvox

项目

项目 研究方向 已实现内容
wavrail 流式语音识别 WAV I/O 与帧化、log-mel 特征、CTC 贪心/前缀束解码、transducer 式帧同步发射、强制对齐、中英文 CER/WER、分块流式管线与延迟核算、合成音频上的可训练声学模型
dialogvox 端到端语音对话 RVQ 神经音频编解码、语音-文本统一 token、紧凑对话 LM 与联合损失、KV-cache 流式生成与有界缓冲、外部语音大模型适配器协议、回合级评测与可复现报告

两个项目都包含命令行工具、可运行示例、契约/黄金测试和 CPU 模型测试,全部离线可复现。

技术栈

Python NumPy PyTorch pytest ruff GitHub Actions

活动

contribution snake

当前关注

  • 流式识别的延迟核算:算法延迟 vs 发射延迟,分块与离线输出的等价性。
  • 语音-文本统一 token 上的联合建模、KV-cache 流式解码与背压语义。
  • 合成数据上的可复现训练链路:种子、锁文件与逐提交检查。

Popular repositories Loading

  1. wavrail wavrail Public

    Streaming speech recognition toolkit: log-mel features, CTC greedy and prefix-beam decoding, transducer-style frame-synchronous emission, forced alignment, Chinese/English CER-WER metrics, and a ch…

    Python 31 235

  2. StarWorkshop StarWorkshop Public

    Profile README

  3. dialogvox dialogvox Public

    End-to-end spoken dialogue toolkit: trainable RVQ neural audio codec, unified speech-text tokens, compact decoder-only LM with joint losses, KV-cache streaming generation, offline-tested adapter pr…

    Python