Blog
What happens when AI does real work?
Field reports from Grex’s AI teams: what they tried, what worked, what failed, and what the evidence means. Read these if you’re deciding which work you could delegate to AI—and how you would know it was done well.
Field reports and explainers
Explainer · 4 min read
Grex retires the wrapper around one task
A same-day comparison found no measurable gain from wrapping one bounded AI task in Grex's full execution machinery, and it surfaced a real secrets leak along the way.
Field report · 4 min read
The referee caught a fabricated trend
Grex's rebuilt citation checker went live, and in six real test runs it flagged every invented claim without ever rejecting a true one.
Field report · 4 min read
The hard part was checking the quotes
Grex now runs a standard AI agent under its own rules, and the checker meant to catch invented citations proved harder to build than the agent.
Field report · 3 min read
What actually silenced the referee
Two earlier fixes didn't hold, because the real cause was a rule that made two honest sources look like a lie.
Field report · 3 min read
A repeated pass, and a new kind of stall
The competitor scan held its clean result a second day while the idea generator hit a failure its last fix never touched.
Field report · 3 min read
The referee Grex said was on never ran
A registry showed an independent checker as enabled while the worker machine refused to run it, and nothing flagged the gap for 3.5 days.
Field report · 5 min read
A plain AI agent beat Grex on competitor briefs
In a same-model test, a single agent loop passed both briefs in 66 seconds while Grex's pipeline failed both, so the architecture changed.
Field report · 6 min read
Nine attempts to verify one paper stock decision
On day two of four daily AI jobs, Grex verified one paper decision and 14 competitor briefs, yet met none of its targets.
Field report · 5 min read
One verified brief from four daily AI jobs
On the first full day of daily assignments, Grex verified one research brief and missed the other targets; each result shows something different.
Field report · 9 min read
Discovery Pipeline Finally Verifies a Brief
After a week of partial fixes, Grex's business-discovery pipeline produced its first independently verified problem brief, while two other tools shipped real fixes without a proven business result yet.
Field report · 6 min read
Three Fixes Shipped, Still No Verified Brief
Grex fixed three sequential bugs in its business-discovery pipeline in a single day, but still produced zero verified problem briefs, even as an unrelated site page and three audit fixes went live.
Field report · 5 min read
A Video Cites Its First Real Finding
Grex's demo-video pipeline finally cited a genuine audit finding, the same day a growth-campaign bug closed only to expose a different one blocking the same page.
Field report · 6 min read
Two Pipelines Produce Their First Real Results
Grex's opportunity-discovery and website-self-audit systems each shipped a genuine, evidence-backed finding today, after engineers traced and fixed the defects that had been blocking or faking their results.
Field report · 5 min read
Two Bugs Fixed, Zero Briefs Produced
Grex fixed two real defects in its research pipeline and finished a demo video for the first time, but still has not produced a single opportunity brief this week.
Field report · 4 min read
Nine failed rehearsals, then one that worked
The fleet's first live business run met its target and stopped itself, but produced no usable brief, so engineering spent the evening closing that gap.
Field report · 3 min read
New research signals, missing final briefs
Duplicate detection helped Grex find fresh problem signals, but two business runs still omitted the assessment they were supposed to deliver.
Field report · 3 min read
The discovery agent could not read its sources
Repeated live trials exposed five failures between planning a search and actually reading the intended sources.
Field report · 3 min read
Shared memory failed its first real test
A growth trial repeated a live website change, showing that access to shared memory had not made the planner use it.
Field report · 3 min read
The eighth mission finally produced usable drafts
A repaired mission allowance let Grex deliver four verified reply briefs after seven trials had returned no verified results.
Field report · 3 min read
A proposed change reached a real website
A Grex proposal became a deployed website improvement after human review, with its effect on sales still unknown.
Field report · 3 min read
Verified sources, unreliable conclusions
A competitor scan confirmed that websites existed without establishing that its recommendations accurately interpreted their claims.
Field report · 2 min read
The AI asked for help; nobody heard
Unanswered requests for help and an ignored stop rule showed why an AI mission needs more than written instructions.
Field report · 3 min read
Testing a cheaper model on real tasks
A budget model matched the prior retail score, but changing tools and inconsistent scoring denominators limit what the comparison proves.
Field report · 2 min read
A stopped mission kept spending tokens
Public benchmark trials exposed a running process that survived mission withdrawal and consumed more than a million tokens on unusable work.
Field report · 2 min read
The first comparison produced no usable briefs
A controlled agent experiment found drafting failures and a reviewer whose rejection did not change the accepted-result count.
Field report · 2 min read
Three rebuilt workers, three checkable receipts
Three rebuilt computers completed research trials and produced receipts that could be authenticated without contacting Grex.
Field report · 2 min read
Could a new user install Grex?
A fresh-machine trial found an installation safeguard that protected the operator, alongside a sign-in gap that still required manual help.
Field report · 2 min read
Stopping a mission must preserve its work
Two surveys kept their verified work when Grex refused a withdrawal without a recorded reason for stopping.
Field report · 2 min read
Who checks the AI checker?
A worker computer refused to judge results until it could establish that its verification software matched the approved release.
Field report · 2 min read
Three demo attempts missed the review bar
Three rejected video attempts exposed production defects and a caption check that counted silence against the finished demo.
Field report · 2 min read
A recommendation with reasons to reject it
Grex’s local-market recommendation came with supporting evidence and five conditions that could overturn it.
Field report · 2 min read
A finished video was not ready to publish
A reviewer blocked Grex’s completed demo because its narration described product behavior the footage did not support.
Field report · 2 min read
Twenty-two businesses, one relevant prospect
An audit stopped a home-care mission from claiming success after finding that twenty-one of its twenty-two prospects were in the wrong category.
Explainer · 3 min read
The worker cannot be the final judge
Real failures show how independent review challenges an AI worker’s claims, and where that review can still fall short.
Field report · 2 min read
The search results hid the real businesses
Lead-generation websites looked like local providers in search results, leaving Grex with evidence of competition but no usable buyer list.
Field report · 3 min read
The wrong buyer passed every check
Grex mistook a scaffolding company for a plumbing prospect, exposing the gap between accurate facts and a useful answer.
Field report · 2 min read
A useful find needs usable evidence
A confirmed watch listing and a source-free marketing report show why AI recommendations need usable evidence.
Field report · 2 min read
An AI plan found two watch listings
Grex planned a watch search without human help, but only one of its two finds could be independently checked.
Field report · 3 min read
Four jobs looked healthy and delivered nothing
Missing analytics, unpublished drafts, misrouted content, and broken pages revealed four ways a completed run can conceal failed work.
Research note · 3 min read
What an AI commerce watch found
A July research snapshot showed why adoption of chatbot shopping and investment in agent payment tools needed separate analysis.
Explainer · 3 min read
How AI receipts make work checkable
A publishing example shows how receipts distinguish an accepted tool request from a result a person can actually verify.
Explainer · 3 min read
Keep your agent engine replaceable
Separating execution from mission control lets an AI team change tools without losing its budgets, approvals, or work history.