Mon 15 Jun 2026 12:00 - 12:20 at Meadows B - PAgE Session 1 Chair(s): Shraddha Barke

Large Language Models (LLMs) are increasingly used for code generation, including complex distributed systems. However, their reliability in stateful and failure-prone environments remains unclear. In this paper, we present a structured evaluation framework for assessing LLM-generated implementations of distributed protocols, including Two-Phase Commit, Ring Election, and Raft. We introduce a simulation environment that injects realistic network faults such as message loss, delay, and duplication. Our results show that while LLMs perform well on simpler protocols and code completion tasks, they struggle with complex consensus algorithms and exhibit inconsistent debugging behavior. We further analyze the relationship between model size and performance, finding that parameter count alone does not predict effectiveness. These findings highlight the need for systematic testing and verification in agentic code generation systems.

Mon 15 Jun

Displayed time zone: Mountain Time (US & Canada) change

10:40 - 12:20
PAgE Session 1PAgE at Meadows B
Chair(s): Shraddha Barke Microsoft Research, Redmond
10:40
20m
Talk
Agentic Code Reasoning
PAgE
Shubham Ugare Meta, Satish Chandra Meta Platforms, Inc.
Pre-print
11:00
20m
Talk
Towards Verified Code Reasoning by LLMs
PAgE
Meghana Aparna Sistla Google DeepMind, Gogul Balakrishnan Google, Patrick Rondon Google, José Pablo Cambronero Google, USA, Michele Tufano Google, Satish Chandra Meta Platforms, Inc.
11:20
20m
Talk
Testing, Credible Compilation, and Verification in the Axon Verified Compiler in Lean and Claude CodeRemote
PAgE
Martin C. Rinard Massachusetts Institute of Technology
11:40
20m
Talk
The Next Frontier for AI-Generated Kernels: Correctness
PAgE
Guido Martínez Microsoft Research, Tyler Sorensen University of California at Santa Cruz
DOI Pre-print
12:00
20m
Talk
Testing LLM-Generated Distributed Protocol Code
PAgE
Ankush Das Boston University, Brendan Coyne Boston University