From Workarounds to Root Causes: Experience Using Agentic Workflows to Debug Complex Browser GPU Compiler Stacks
We report on porting llama.cpp, a widely used LLM inference engine, to the browser via WebGPU. Failures in this setting are hard to explain: one wrong value can pass through shader code, browser runtime, compiler rewrites, native GPU tools, and GPU execution before becoming visible. llama.cpp’s debugging tools localize many problems, but the remaining cases can take days or weeks of expert effort across an unfamiliar stack.
We built a context-managed agentic workflow in which scoped agents use real browser and compiler tools, communicate through file-based records, and rely on reviewer prompts to decide whether evidence is sufficient. In our case study, a quantized byte-unpacking shader returned zeros in Chrome on Apple Silicon. The workflow found a bit-equivalent source workaround in ~2 h active agent time (~3 h wall-clock with one token-quota refresh). Because the workaround did not explain the platform bug, we changed the reviewer prompt to require mechanism-level evidence. A local-array shape-and-size trigger matrix then took ~12 h active across six sessions (~3.5 d wall-clock paced by daily quota refreshes), and a subsequent 47-min Metal-compiler session produced a byte-identical intermediate-output comparison, a standalone Metal reproducer, and a GPU loop-unroller root cause. The experience suggests that re-steerable workflows can shift complex browser/GPU debugging from weeks-scale manual digging toward lower-overhead monitoring, steering, and artifact validation.
Mon 15 JunDisplayed time zone: Mountain Time (US & Canada) change
13:40 - 15:20 | |||
13:40 20mTalk | Combining Agentic AI and Lightweight Formal Methods To Find Bugs in a Production Hypervisor PAgE Hiroyuki Katsura University of Cambridge, Kayvan Memarian University of Cambridge, Peter Sewell University of Cambridge | ||
14:00 20mTalk | From Workarounds to Root Causes: Experience Using Agentic Workflows to Debug Complex Browser GPU Compiler Stacks PAgE Abhijit Ramesh UC Santa Cruz, Reese Levine University of California at Santa Cruz, Tyler Sorensen University of California at Santa Cruz | ||
14:20 20mTalk | Event-based Design Abstractions for Agent HarnessesRecorded PAgE McCoy Becker , Matin Ghavami Massachusetts Institute of Technology, Fabian Zaiser Massachusetts Institute of Technology, Timothy J. O'Donnell McGill University; Mila – Quebec AI Institute; CHI FRO; Canada CIFAR AI Chai, Mila, Martin C. Rinard Massachusetts Institute of Technology, Joshua B. Tenenbaum Massachusetts Institute of Technology, Vikash Mansinghka Massachusetts Institute of Technology | ||
14:40 20mTalk | Lumos: Let there be Language Model System CertificationRemote PAgE Isha Chaudhary , Vedaant Jain , Prineet Parhar , Kavya Sachdeva , Avaljot Singh University of Illinois Urbana-Champaign, Sayan Ranu , Gagandeep Singh University of Illinois Urbana-Champaign Pre-print | ||
15:00 20mTalk | Lazy Validation and Self-Healing for Agentic ProgramsRemote PAgE Theodoros Tsampouris Aristotle University of Thessaloniki, Eleftherios Ioannidis Microsoft Research, Andreas Symeonidis Aristotle University of Thessaloniki | ||