AI QA automation for continuous, quick iterations
A pipeline that turns a narrated screen recording of implementation QA into a reviewable list of timestamped issues, then, once approved, a filed GitHub issue with a screenshot attached.
Role :
Design lead
Product Strategy
Runtime & Canvas:
Claude code
Gemini
Timeline :
May 2026
Skills :
Product Design
AI automation
The problem
The spreadsheet was the bottleneck, not the testing
Once a design gets implemented, it's a task to QA the build against it, page by page, checking that what shipped actually matches what was designed. The part that hurt was everything after. Every mismatch I found had to be typed into a row on a shared sheet.
Page, issue, current behaviour, expected behaviour, testing criteria, plus two screenshots : one of how it looked in the build, one of how it was supposed to look in the design.
Where the idea came from
I wasn't looking for a QA tool
I was already using Gemini AI Studio elsewhere in my job. It's genuinely good at analysing a video and telling you what's in it, so I'd been feeding it user research recordings to pull out themes faster than re-watching everything myself.
One afternoon, while working on QA for a design feature we launched and I didnt want to make the spreadsheet, it clicked. I screen recorded the going through the feature that was shipped and just mentioned what works and whats off.
Testing was never the slow part. Translating onto the spreadsheet someone else could act on and track was.
So, three decisions
Instead of a better template
The analyze skill
A recording becomes a reviewable list, not a filed issue
The dual product strategy delivered measurable business outcomes - from increased engagement to qualified pipeline growth. Bounce rate of the landing page also reduced by 11%.
Experience-led strategy improved website conversion and reduced bounce rates
Partnership approach enabled best-in-class voice A without building everything in-house
Dual product strategy solved both technology and adoption challenges for Indian market
The pipeline
Four stages, one thread of context
Each stage hands the next a clean, structured artifact, not a raw video and a hope.
Proof
What the pipeline actually checks
The outcome
Github issue with the description, details, screenshots, etc.
Results
What changed for the team
Not a paragraph of prose, one entry out of a real analysis pass, same six fields every time, watched and heard straight off the video rather than a transcript.
Learnings
The hard part was never finding the bug
The bottleneck was never the testing. I could already find the bugs, that part didn't need fixing, and Relay didn't change how I test at all. It changed what happens after I notice something: no more stopping mid-test to log it into Excel and go hunting for a screenshot. That interruption, not the testing itself, turned out to be most of the work.
Trust isn't something you can argue someone into, either. The first few tickets a bot filed got questioned. What actually changed people's minds was the review step: once every issue was checked before filing, and kept being right pass after pass, "a bot filed this" stopped being a red flag.
If I built something like this again, I'd start from the smallest lesson here: logging live didn't just cost the time it took to type an issue up, it broke my testing flow. Getting back into a testing headspace after switching over to Excel was its own tax, one I didn't really notice until recording removed it.
"Most bugs don't die from being missed. They die in the gap between noticing and fixing."




