Xuzheng He
- RWKV: Reinventing RNNs for the Transformer Era
- StableMask: Refining Causal Masking in Decoder-only Transformer
- How Do Humans Write Code? Large Models Do It the Same Way Too
- Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
- Interpretable Contrastive Monte Carlo Tree Search Reasoning