Yukyung Lee
  • Home
  • News
  • Publications
  • Teaching
ESC

Searching...

No results found

โ†‘โ†“ Navigate โ†ต Select
Powered by Hugo Blox
  • Projects
  • Experience
  • News
    • ๐ŸŽ‰ "Harbor Adapters and Harbor-Index" has been accepted to NeurIPS 2026 (E&D)
    • ๐ŸŽ‰ Three papers accepted at EMNLP 2026 Main conference!
    • ๐Ÿ’ก Check out our math reasoning paper "Cliff Token"
    • ๐Ÿ’ก Check out our implicit CoT paper "CIRF"
    • ๐ŸŽ‰ "Can Structural Cues Save LLMs?" has been accepted to KDD 2026!
    • ๐ŸŽ‰ RExBench has been accepted to ACL 2026!
    • ๐Ÿ‘ฉ๐Ÿปโ€๐Ÿซ Invited talks at SNU (Jan 2nd), HYU (Jan 8th), and KU (Jan 9th)
    • ๐Ÿ‘ฉ๐Ÿปโ€๐Ÿซ Gave an invited talk at Stanford/UW (RExBench ๐Ÿฆ–) with Nicholas
    • ๐ŸŽ‰ CheckEval has been accepted to EMNLP 2025!
    • ๐Ÿ‘ฉ๐Ÿปโ€๐Ÿซ Gave an invitied talk at Korea University (RExBench ๐Ÿฆ–)
    • ๐Ÿ’ก Check out our RExBench paper
    • ๐ŸŽ‰ WritingPath has been accepted to NAACL 2025 (Industry Track)!
    • ๐ŸŽ‰ Paper accepted to NeurIPS 2024
    • ๐Ÿ‘ฉ๐Ÿปโ€๐Ÿซ Gave an invitied talk at Korea University
    • ๐Ÿ“ฃ I have started my Postdoc at Boston University @ tinlab
    • ๐Ÿ‘ฉ๐Ÿปโ€๐Ÿซ Gave an invited talk at SK Telecom
  • Publications
    • Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation
    • AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not
    • Decoupled Semantic Compression: KV Cache Eviction with Semantic Localization and Selection
    • LLMs Learn Better In-Context from Rules than from Examples
    • PGMem: Tightly Coupled Persona-Memory Graph for Lifelong Personalized Agents
    • Position: Artificial Hivemind Evidence Does Not Support Societal Homogenization Claims
    • Cliff Tokens: Analyzing Failure Trigger Tokens in LLM Mathematical Reasoning
    • CIRF: Tokenizing Chain-of-Thoughts into Reusable Functional Units for Efficient Latent Reasoning in Large Language Models
    • Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams
    • RExBench: Can coding agents autonomously implement AI research extensions?
    • CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
    • Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models
    • A Gradient Accumulation Method for Dense Retriever under Memory Constraint
    • DSTEA: Improving Dialogue State Tracking via Entity Adaptive Pre-training
    • RAPID: Training-free Retrieval-based Log Anomaly Detection with PLM considering Token-level information
    • LAnoBERT: System log anomaly detection based on BERT masked language model
    • Painsight: An Extendable Opinion Mining Framework for Detecting Pain Points Based on Online Customer Reviews
    • Oh My Mistake!: Toward Realistic Dialogue State Tracking including Turnback Utterances
    • Mismatch between Multi-turn Dialogue and its Evaluation Metric in Dialogue State Tracking
    • Multi^2OIE: Multilingual Open Information Extraction Based on Multi-Head Attention with BERT
    • Drone Surveillance System Considering Dynamic POIs
  • Recent & Upcoming Talks
    • Example Talk
  • Teaching Experience

๐ŸŽ‰ RExBench has been accepted to ACL 2026!

Apr 6, 2026 ยท 1 min read

RExBench, Can coding agents autonomously implement AI research extensions?

Last updated on Jun 25, 2026
Yukyung Lee
Authors
Yukyung Lee
Postdoctoral Associate

← ๐ŸŽ‰ "Can Structural Cues Save LLMs?" has been accepted to KDD 2026! May 15, 2026
๐Ÿ‘ฉ๐Ÿปโ€๐Ÿซ Invited talks at SNU (Jan 2nd), HYU (Jan 8th), and KU (Jan 9th) Jan 2, 2026 →

ยฉ 2026 Yukyung Lee. This work is licensed under CC BY NC ND 4.0

Made with Hugo Blox. Create your site โ†’