Mon 15 Jun 2026 14:00 - 14:20 at Meadows B - PAgE Session 2

We report on porting llama.cpp, a widely used LLM inference engine, to the browser via WebGPU. Failures in this setting are hard to explain: one wrong value can pass through shader code, browser runtime, compiler rewrites, native GPU tools, and GPU execution before becoming visible. llama.cpp’s debugging tools localize many problems, but the remaining cases can take days or weeks of expert effort across an unfamiliar stack.

We built a context-managed agentic workflow in which scoped agents use real browser and compiler tools, communicate through file-based records, and rely on reviewer prompts to decide whether evidence is sufficient. In our case study, a quantized byte-unpacking shader returned zeros in Chrome on Apple Silicon. The workflow found a bit-equivalent source workaround in ~2 h active agent time (~3 h wall-clock with one token-quota refresh). Because the workaround did not explain the platform bug, we changed the reviewer prompt to require mechanism-level evidence. A local-array shape-and-size trigger matrix then took ~12 h active across six sessions (~3.5 d wall-clock paced by daily quota refreshes), and a subsequent 47-min Metal-compiler session produced a byte-identical intermediate-output comparison, a standalone Metal reproducer, and a GPU loop-unroller root cause. The experience suggests that re-steerable workflows can shift complex browser/GPU debugging from weeks-scale manual digging toward lower-overhead monitoring, steering, and artifact validation.

Mon 15 Jun

Displayed time zone: Mountain Time (US & Canada) change

13:40 - 15:20
PAgE Session 2PAgE at Meadows B
13:40
20m
Talk
Combining Agentic AI and Lightweight Formal Methods To Find Bugs in a Production Hypervisor
PAgE
Hiroyuki Katsura University of Cambridge, Kayvan Memarian University of Cambridge, Peter Sewell University of Cambridge
14:00
20m
Talk
From Workarounds to Root Causes: Experience Using Agentic Workflows to Debug Complex Browser GPU Compiler Stacks
PAgE
Abhijit Ramesh UC Santa Cruz, Reese Levine University of California at Santa Cruz, Tyler Sorensen University of California at Santa Cruz
14:20
20m
Talk
Event-based Design Abstractions for Agent HarnessesRecorded
PAgE
McCoy Becker , Matin Ghavami Massachusetts Institute of Technology, Fabian Zaiser Massachusetts Institute of Technology, Timothy J. O'Donnell McGill University; Mila – Quebec AI Institute; CHI FRO; Canada CIFAR AI Chai, Mila, Martin C. Rinard Massachusetts Institute of Technology, Joshua B. Tenenbaum Massachusetts Institute of Technology, Vikash Mansinghka Massachusetts Institute of Technology
14:40
20m
Talk
Lumos: Let there be Language Model System CertificationRemote
PAgE
Isha Chaudhary , Vedaant Jain , Prineet Parhar , Kavya Sachdeva , Avaljot Singh University of Illinois Urbana-Champaign, Sayan Ranu , Gagandeep Singh University of Illinois Urbana-Champaign
Pre-print
15:00
20m
Talk
Lazy Validation and Self-Healing for Agentic ProgramsRemote
PAgE
Theodoros Tsampouris Aristotle University of Thessaloniki, Eleftherios Ioannidis Microsoft Research, Andreas Symeonidis Aristotle University of Thessaloniki