🎉 Three papers accepted at EMNLP 2026 Main conference!
➊ AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not ➋ PGMem: Tightly Coupled Persona–Memory Graph for Lifelong Personalized Agents and ➌ LLMs Learn Better …
➊ AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not ➋ PGMem: Tightly Coupled Persona–Memory Graph for Lifelong Personalized Agents and ➌ LLMs Learn Better …
'PGMem: Tightly Coupled Persona–Memory Graph for Lifelong Personalized Agents’ is now available on [arXiv](https://arxiv.org/abs/2608.01708)! Work with Wonjun Choi, Yerim Kim and …
'Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning' now available on [arXiv](https://arxiv.org/abs/2606.25524)! Work with Jaeyong Ko and Pilsung …
'CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models' now available on …
Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams. Work with Yebin Lim, Woojun Jung, Wonjun Choi and Susik Yoon 🐯
RExBench, Can coding agents autonomously implement AI research extensions?
Topic - Towards a Science of Evaluation for Language Model
Topic - Can coding agents autonomouslyimplement AI research extensions?
CheckEval, A reliable LLM-as-a-Judge framework for evaluating text generation using checklists. Thanks to my collaborators ❤️
Topic - Can coding agents autonomously implement AI research extensions?