Experiments · bench notes
Built, not theorized.
Four entries. One is live on the web,
one runs inside a company every week,
two began at hackathons.
Not side projects. Working products, each testing the same idea: AI should know when to ask a human.
TruthLens ↗
Reads insurance denial letters, finds the mistake, writes the appeal. Built in 4 days for the Google DeepMind Gemini 3 hackathon. The idea: an AI can carry a fight most people abandon, if it shows its reasoning well enough for a person to sign their name to it.
Slide Surgeon ↗
Upload a messy slide, get a pitch-ready one in about a minute. Built entirely with prompting for the Nano Banana hackathon, with no hand-written pipeline code. The idea: design judgment can be written down precisely enough that a model can execute it.
Design Review Agents
Five AI agents that cut design reviews at Automation Anywhere from a week to hours. Eight designers use them, 20+ reviews a month. Each agent reviews one thing (accessibility, consistency, content, flows, edge cases) and flags what needs a human eye instead of pretending to be one.
ML Before the Roadmap
A prediction prototype at a Siemens hackathon that proved the case for ML 18 months before it reached the official roadmap. The oldest entry here, and the pattern the rest follow: build the argument, don't present it.
Got a hypothesis worth a weekend? Bring it.