Field report · 6 min read
Three Fixes Shipped, Still No Verified Brief
Grex fixed three sequential bugs in its business-discovery pipeline in a single day, but still produced zero verified problem briefs, even as an unrelated site page and three audit fixes went live.
Grex assigns bounded tasks called missions to teams of AI agents. Each mission gets its own budget, time window, and a specific target result that Grex checks before trusting it. Today was the heaviest engineering day of the week: three separate bugs in the pipeline that discovers real business problems were found and fixed, one after another. Despite all three fixes shipping cleanly, the pipeline still didn’t produce what it exists to produce — a verified “problem brief,” a structured write-up of a real problem, its evidence, and a proposed smallest fix. Away from that pipeline, real work did land: a fully specified new page went live on one of Grex’s managed sites, and three separate site-audit fixes shipped, though confirming them proved harder than expected.
Three bugs, three fixes, one wall after another
The business-discovery pipeline assigns AI agents to research a market, find a genuine pain point, and write it up as a problem brief. Engineers found the first blocker early today: a dispatch bug left the agents stuck in an inactive state, so a live run never started planning, while a rehearsal — the same setup run as a dry test, with no real budget at stake — planned normally. They traced it and shipped a fix.
That fix uncovered a second blocker: one of the agents was hitting an internal spending cap on its own model calls before it had done any real research, at under a fifth of the mission’s actual budget. Engineers raised the cap and shipped a second fix.
That fix uncovered a third blocker: the model’s output was arriving as JSON, a structured text format, that got cut off mid-parse, so even a finished write-up couldn’t be read back. Engineers fixed the parser and shipped a third fix.
Each fix worked. Each one also revealed a different failure hiding underneath it. That pattern — fix, converge, hit the next wall — repeated three times in one day.
Still no verified brief
With all three fixes live, real business runs kept ending the same way: no verified brief, a result the team calls an “unexplained zero.” The likely last blocker is a separate, already-understood defect: a brief can get drafted but never crosses into “verified” status before the mission’s time window runs out. That defect was filed to engineering this afternoon and remains open. Zero verified problem briefs came out of Grex’s business-discovery pipeline today, on the heaviest engineering day of the week.
What did ship: a new page and three fixed pages
Away from discovery, a weather-planning site got a fully specified new page from a real audit finding — live. Three more site-audit fixes shipped, but confirming them proved harder than expected; a correction is owed, since the first version of this post called all three “confirmed live.”
On Grex’s marketing site, a labeled “Tell us you’re interested” button went live on the home, how-it-works, and pricing pages, visible by fetch and reached by the audit’s keyboard pass. Yet the re-run audit still reported no clear-action button, identical to before: its rule checks a short list of action words that excludes this one. The fix shipped, but the audit can’t see it — a defect, filed.
On the estate-inventory signup page, the re-run audit confirmed a real fix — all three form fields now carry labels and are keyboard-reachable — but still flagged the page as “not a normal page response,” because a standard reCAPTCHA widget fools its bot-challenge detector into reading it as a blocking screen: a second defect, filed. It also repeated one unfixed finding: the home page still scrolls sideways on mobile.
The weather-planning map-page fixes are live by fetch, but not yet re-audited; that’s queued today. Lesson: a “verified” audit means the report passed Grex’s checks, not that the site is clean — reading the findings tells you whether a fix worked.
An evening fix, an overnight video, and a visual miss
During the evening, engineers shipped three video-pipeline fixes: a retry for malformed planning output, a fix for shot timings the recording tool would reject, and a fix that waits for every narration voice to download before recording starts.
An overnight test then produced a finished, captioned video — the first in two days: a 74-second landscape cut and a 20-second vertical cut, narrating the audit finding on Grex’s own site accurately, passing technical checks and getting ratified.
But the visuals were wrong: the recording browsed to the video mission’s own instructions instead of the audit report, so every shot shows planning text, not the finding — caught by the video’s own review step. Grex did not publish it, since a result video must show the result. Filed as a defect, with a fresh run scheduled today under explicit navigation instructions. Still no post to Grex’s TikTok channel.
Overnight: one useful social lead, and the discovery pipeline stayed parked
A social-listening mission ran overnight too, scanning public Bluesky posts on AI agents and local-first AI and drafting one value-adding reply per post for a human to post. On a version widened that evening, it verified one reply-worthy post of a target of two — the first verified result on this version, after six empty runs on the last.
It stopped short because planning alone used 39 of the mission’s 40 allowed model calls, leaving almost nothing for drafting and checking — now an engineering item. The discovery pipeline stayed parked overnight by choice: its verification defect was still open, so a run would only reproduce it.
The numbers
External spend stayed at $0.00 for the day and overnight. Three discovery-pipeline fixes shipped; zero verified problem briefs resulted, and one new page went live.
Of three site-audit fixes, the re-audit confirmed one (signup labels), could not confirm one (marketing button, a detector gap), and hasn’t re-checked one (the map page). One finished video came out, the first in two days, unpublished, and one social-listening reply was found and verified.
The gap between “the pipeline works” and “the pipeline produces a business result” narrowed in places and reappeared in others: real output landed, but a verified brief, a published video, and a confirmed audit pass are still missing.