💡 Check out our math reasoning paper "Cliff Token"
'Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning' now available on [arXiv](https://arxiv.org/abs/2606.25524)! Work with Jaeyong Ko and Pilsung …
'Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning' now available on [arXiv](https://arxiv.org/abs/2606.25524)! Work with Jaeyong Ko and Pilsung …
'CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models' now available on …
Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams. Work with Yebin Lim, Woojun Jung, Wonjun Choi and Susik Yoon 🐯
RExBench, Can coding agents autonomously implement AI research extensions?
Please consider submitting your work :)
Topic - Towards a Science of Evaluation for Language Model
Topic - Can coding agents autonomouslyimplement AI research extensions?
CheckEval, A reliable LLM-as-a-Judge framework for evaluating text generation using checklists. Thanks to my collaborators ❤️
Topic - Can coding agents autonomously implement AI research extensions?
Can coding agents autonomously implement AI research extensions? Our [RExBench](https://arxiv.org/abs/2506.22598) is now available on arXiv !