Publications (* equal contribution)

2026

  1. Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
    Yerim Oh , Young-Jun Lee, Jaewoo Ahn, Gunhee Kim, and Dongyeop Kang
    Under Review, 2026
  2. Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
    Under Review, 2026
    Also presented at Workshop on Agent Behavior @ COLM 2026, EAD Workshop @ ECCV 2026, LASS Workshop @ CIKM 2026
  3. Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games
    In ICLR, 2026
    Also presented at Wordplay Workshop @ EMNLP 2025 (Outstanding, Lightning Talk)

2025

  1. FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
    In EMNLP, 2025
    Also presented at Wordplay Workshop @ EMNLP 2025 (Spotlight, Lightning Talk)
  2. ChartCap: Mitigating Hallucination of Dense Chart Captioning
    Junyoung Lim, Jaewoo Ahn, and Gunhee Kim
    In ICCV, 2025 Highlight
  3. Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
    Jaewoo Ahn* , Heeseung Yun*, Dayoon Ko, and Gunhee Kim
    In ACL, 2025
  4. Is a Peeled Apple Still Red? Evaluating LLMs’ Ability for Conceptual Combination with Property Type
    In NAACL, 2025 Oral

2024

  1. TimeChara: Evaluating Point-in-Time Character Hallucination of Role-Playing Large Language Models
    In ACL Findings, 2024
    Also presented at NLRSE Workshop @ ACL 2024
  2. Who Wrote this Code? Watermarking for Code Generation
    In ACL, 2024

2023

  1. mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with Images
    In EMNLP, 2023
  2. MPCHAT: Towards Multimodal Persona-Grounded Conversation
    Jaewoo Ahn, Yeda Song, Sangdoo Yun, and Gunhee Kim
    In ACL, 2023

2020

  1. Sequential Latent Knowledge Selection for Knowledge-Grounded Dialogue
    Byeongchang Kim, Jaewoo Ahn, and Gunhee Kim
    In ICLR, 2020 Spotlight