Skip to content
View lzwhehe's full-sized avatar
🎯
Focusing
🎯
Focusing
  • JAIST

Highlights

  • Pro

Block or report lzwhehe

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
lzwhehe/README.md
Liu Zhuowen — AI / Agent Security Researcher

中文 English 日本語

Email


你好,我是刘卓文,目前在北陆先端科学技术大学院大学(JAIST)先端科学技术专攻 · 网络安全实验室攻读硕士二年级。

我的研究聚焦安全运营场景下的 Agent 安全,致力于让 LLM 智能体安全、可靠地落地于实际的安全运营工作。同时,我也投身于中国非物质文化遗产保护,专注于用前沿 AI 技术助力传统文化的数字化保存与传承。

🔬 Research

Now: premature verdicts under forged evidence

HESP

Passing the Test You Trained On

Safeguarding Intangible Heritage

每个项目一句话
  • 进行中 · 伪造证据下的过早决策:测量 0.5B–72B 开源模型(Qwen2.5、Qwen3、Llama-3.1、Meta-SecAlign)约 1 万个回合;能结案的大模型在伪造日志下把 36 个攻击中的 30–36 个判为良性,Qwen2.5-72B 平均只查 1.1 个数据源就采信。正在预注册“印证感知”训练:训练中不使用任何攻击样本。
  • HESP(第一作者,arXiv:2609.33446):控制器接管“查什么”与“何时停”,模型只负责提议;计数似然表使 Qwen2.5-7B 完成率 0.125 → 1.000,控制器结案使 Llama-3.1-8B 0 → 0.917(4 项预注册研究、7,272 个审计回合)。
  • Passing the Test You Trained On(论文草稿):15 个提示词注入检测器的排名在基准之间几乎不迁移(Kendall τ 0.01–0.31),跑分主要反映基准与训练数据有多接近。
  • 库淑兰彩贴剪纸生成(共同一作,npj Heritage Science 审稿中):从作品自动恢复纸色谱与剪切线,微调 Qwen-Image-Edit,线条召回 0.97。

🧰 Skills

Web 安全  Burp Suite SQLMap Nmap Xray OWASP Top 10

Agent 工具  Claude Code Codex OpenCode WorkBuddy Kimi Code

开发能力  Vue Node.js vLLM Ollama LoRA Hyperledger Fabric

编程语言  Python JavaScript TypeScript C

Popular repositories Loading

  1. HESP HESP Public

    让 7B 模型胜任告警分诊 | Making 7B models capable of alert triage

    TeX 1

  2. i483-source_code_LIU i483-source_code_LIU Public

    i483's source code

    Python

  3. Progress-Todo Progress-Todo Public

    todo- show progress

    TypeScript

  4. kushulan-papercut-blora kushulan-papercut-blora Public

    Ku Shulan papercut research archive: datasets, SDXL B-LoRA checkpoints, generated results, and verified migration tools.

    Python

  5. benign-instruction-bench benign-instruction-bench Public

    Re-evaluating prompt-injection detectors on LLM agent tool outputs (paper draft, scripts, scores)

    Python

  6. lzwhehe lzwhehe Public

    Profile README