| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Original HTTPS Page] |
Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.
You must be logged in to block users.
Contact GitHub support about this user’s behavior. Learn more about reporting abuse.
Report abuseWebsite · Publications · Google Scholar · X · CV · hzy2210@gmail.com
I publish papers as Zhuoyuan Hao and use Larry Hao professionally; both names refer to the same person.
CS undergrad at HITSZ, finishing my thesis two semesters early. I work on LLM reasoning and RL — why models reason the way they do, and where RL training quietly goes wrong — with Jing Li (HITSZ) and Xiaozhi Wang (Tsinghua). In between, I intern in industry and ship the occasional product.
Echo of Prompt — first author, ICLR 2026. Reasoning models almost always restate the question to themselves before they start reasoning. Everyone filed this under "SFT artifact." It's more than that: the echo re-anchors attention and keeps a long chain of thought from drifting, and models pay for it when it's missing. I show the probabilistic cost, trace the information flow, and turn the effect into a prompt-time trick that beats baseline under a fixed token budget. → evidence page · code · OpenReview
Reward Hacking in Rubric-based RL (CHERRL) — co-first author. When an LLM judge hands out the reward, the policy can learn to exploit the judge's blind spots instead of actually improving. Real hacking is covert and tangled with several biases at once, so we built a controllable sandbox: inject a known bias, reproduce the hack cleanly, and pin down the exact step it starts. Then an agent reads the training logs and flags that onset on its own. → project summary · code · arXiv
87 merged pull requests across 17 repositories. The ones with some weight behind them: Mole (60k★, 3 merged), CodexBar (19k★, 22 merged), Sub-Store (10k★, 2 merged), mlx-audio (7.6k★, 2 merged), clash-for-linux (5.7k★, 1 merged) and tectonic (5k★, 2 merged). Still open in cc-switch (121k★) and lm-evaluation-harness (13k★).
LLM reasoning, RL, agents — happy to talk, happier to build. Got a research idea or a prototype that needs to ship? hzy2210@gmail.com
ICLR 2026 code for Echoes as Anchors: Echo-of-Prompt, attention refocusing, and probabilistic analysis of LLM reasoning.
CHERRL: A Controllable Hacking Environment for Rubric-Based Reinforcement Learning
Python 12
Processed / Cleaned Data for Paper Copilot
My first ML hands-on project, Top 0.7% Kaggle House Prices regression solution using XGBoost, Optuna, and practical feature engineering.
Python 1
Show usage stats for OpenAI Codex and Claude Code, without having to login.
| Back | FazBrowse Home | New Git URL |