Blog

What happens when AI does real work?

Field reports from Grex’s AI teams: what they tried, what worked, what failed, and what the evidence means. Read these if you’re deciding which work you could delegate to AI—and how you would know it was done well.

Field reports and explainers

  • Explainer · 4 min read

    Grex retires the wrapper around one task

    A same-day comparison found no measurable gain from wrapping one bounded AI task in Grex's full execution machinery, and it surfaced a real secrets leak along the way.

  • Field report · 4 min read

    The referee caught a fabricated trend

    Grex's rebuilt citation checker went live, and in six real test runs it flagged every invented claim without ever rejecting a true one.

  • Field report · 4 min read

    The hard part was checking the quotes

    Grex now runs a standard AI agent under its own rules, and the checker meant to catch invented citations proved harder to build than the agent.

  • Field report · 3 min read

    What actually silenced the referee

    Two earlier fixes didn't hold, because the real cause was a rule that made two honest sources look like a lie.

  • Field report · 3 min read

    A repeated pass, and a new kind of stall

    The competitor scan held its clean result a second day while the idea generator hit a failure its last fix never touched.

  • Field report · 3 min read

    The referee Grex said was on never ran

    A registry showed an independent checker as enabled while the worker machine refused to run it, and nothing flagged the gap for 3.5 days.

  • Field report · 5 min read

    A plain AI agent beat Grex on competitor briefs

    In a same-model test, a single agent loop passed both briefs in 66 seconds while Grex's pipeline failed both, so the architecture changed.

  • Field report · 6 min read

    Nine attempts to verify one paper stock decision

    On day two of four daily AI jobs, Grex verified one paper decision and 14 competitor briefs, yet met none of its targets.

  • Field report · 5 min read

    One verified brief from four daily AI jobs

    On the first full day of daily assignments, Grex verified one research brief and missed the other targets; each result shows something different.

  • Field report · 9 min read

    Discovery Pipeline Finally Verifies a Brief

    After a week of partial fixes, Grex's business-discovery pipeline produced its first independently verified problem brief, while two other tools shipped real fixes without a proven business result yet.

  • Field report · 6 min read

    Three Fixes Shipped, Still No Verified Brief

    Grex fixed three sequential bugs in its business-discovery pipeline in a single day, but still produced zero verified problem briefs, even as an unrelated site page and three audit fixes went live.

  • Field report · 5 min read

    A Video Cites Its First Real Finding

    Grex's demo-video pipeline finally cited a genuine audit finding, the same day a growth-campaign bug closed only to expose a different one blocking the same page.

  • Field report · 6 min read

    Two Pipelines Produce Their First Real Results

    Grex's opportunity-discovery and website-self-audit systems each shipped a genuine, evidence-backed finding today, after engineers traced and fixed the defects that had been blocking or faking their results.

  • Field report · 5 min read

    Two Bugs Fixed, Zero Briefs Produced

    Grex fixed two real defects in its research pipeline and finished a demo video for the first time, but still has not produced a single opportunity brief this week.

  • Field report · 4 min read

    Nine failed rehearsals, then one that worked

    The fleet's first live business run met its target and stopped itself, but produced no usable brief, so engineering spent the evening closing that gap.

  • Field report · 3 min read

    New research signals, missing final briefs

    Duplicate detection helped Grex find fresh problem signals, but two business runs still omitted the assessment they were supposed to deliver.

  • Field report · 3 min read

    The discovery agent could not read its sources

    Repeated live trials exposed five failures between planning a search and actually reading the intended sources.

  • Field report · 3 min read

    Shared memory failed its first real test

    A growth trial repeated a live website change, showing that access to shared memory had not made the planner use it.

  • Field report · 3 min read

    The eighth mission finally produced usable drafts

    A repaired mission allowance let Grex deliver four verified reply briefs after seven trials had returned no verified results.

  • Field report · 3 min read

    A proposed change reached a real website

    A Grex proposal became a deployed website improvement after human review, with its effect on sales still unknown.

  • Field report · 3 min read

    Verified sources, unreliable conclusions

    A competitor scan confirmed that websites existed without establishing that its recommendations accurately interpreted their claims.

  • Field report · 2 min read

    The AI asked for help; nobody heard

    Unanswered requests for help and an ignored stop rule showed why an AI mission needs more than written instructions.

  • Field report · 3 min read

    Testing a cheaper model on real tasks

    A budget model matched the prior retail score, but changing tools and inconsistent scoring denominators limit what the comparison proves.

  • Field report · 2 min read

    A stopped mission kept spending tokens

    Public benchmark trials exposed a running process that survived mission withdrawal and consumed more than a million tokens on unusable work.

  • Field report · 2 min read

    The first comparison produced no usable briefs

    A controlled agent experiment found drafting failures and a reviewer whose rejection did not change the accepted-result count.

  • Field report · 2 min read

    Three rebuilt workers, three checkable receipts

    Three rebuilt computers completed research trials and produced receipts that could be authenticated without contacting Grex.

  • Field report · 2 min read

    Could a new user install Grex?

    A fresh-machine trial found an installation safeguard that protected the operator, alongside a sign-in gap that still required manual help.

  • Field report · 2 min read

    Stopping a mission must preserve its work

    Two surveys kept their verified work when Grex refused a withdrawal without a recorded reason for stopping.

  • Field report · 2 min read

    Who checks the AI checker?

    A worker computer refused to judge results until it could establish that its verification software matched the approved release.

  • Field report · 2 min read

    Three demo attempts missed the review bar

    Three rejected video attempts exposed production defects and a caption check that counted silence against the finished demo.

  • Field report · 2 min read

    A recommendation with reasons to reject it

    Grex’s local-market recommendation came with supporting evidence and five conditions that could overturn it.

  • Field report · 2 min read

    A finished video was not ready to publish

    A reviewer blocked Grex’s completed demo because its narration described product behavior the footage did not support.

  • Field report · 2 min read

    Twenty-two businesses, one relevant prospect

    An audit stopped a home-care mission from claiming success after finding that twenty-one of its twenty-two prospects were in the wrong category.

  • Explainer · 3 min read

    The worker cannot be the final judge

    Real failures show how independent review challenges an AI worker’s claims, and where that review can still fall short.

  • Field report · 2 min read

    The search results hid the real businesses

    Lead-generation websites looked like local providers in search results, leaving Grex with evidence of competition but no usable buyer list.

  • Field report · 3 min read

    The wrong buyer passed every check

    Grex mistook a scaffolding company for a plumbing prospect, exposing the gap between accurate facts and a useful answer.

  • Field report · 2 min read

    A useful find needs usable evidence

    A confirmed watch listing and a source-free marketing report show why AI recommendations need usable evidence.

  • Field report · 2 min read

    An AI plan found two watch listings

    Grex planned a watch search without human help, but only one of its two finds could be independently checked.

  • Field report · 3 min read

    Four jobs looked healthy and delivered nothing

    Missing analytics, unpublished drafts, misrouted content, and broken pages revealed four ways a completed run can conceal failed work.

  • Research note · 3 min read

    What an AI commerce watch found

    A July research snapshot showed why adoption of chatbot shopping and investment in agent payment tools needed separate analysis.

  • Explainer · 3 min read

    How AI receipts make work checkable

    A publishing example shows how receipts distinguish an accepted tool request from a result a person can actually verify.

  • Explainer · 3 min read

    Keep your agent engine replaceable

    Separating execution from mission control lets an AI team change tools without losing its budgets, approvals, or work history.