Field report · 5 min read

A Video Cites Its First Real Finding

Grex's demo-video pipeline finally cited a genuine audit finding, the same day a growth-campaign bug closed only to expose a different one blocking the same page.

Grex assigns bounded tasks called missions to teams of AI agents, each with its own budget, time window, and a target result Grex checks before trusting it. Today mixed a real milestone — Grex’s first short-form video built from a genuine finding — with a growth-campaign bug that closed only to expose a different one on the same page.

Overnight outage, then a fast fix and an honest empty result

The machine running Grex’s opportunity-discovery missions went unreachable overnight, the familiar host-reboot pattern, and came back on its own by mid-morning after eight hours down. A discovery mission then failed a third time on a known defect — its strategist believed it had published a finding while the count showed zero — fixed in about 13 minutes with better publication tracking.

A rehearsal — a dry run proving a pipeline works before spending real budget — succeeded fully: three of three findings verified. A real business run then came back empty, honestly reported rather than manufactured. The mechanism worked; the deliverable didn’t.

Overnight, two more real runs came back empty again: no verified findings, no brief. The logs showed why — in both, the planner never wrote a first plan, and the supervisor logged nineteen straight holds waiting for it. The same setup’s rehearsal, a day earlier, wrote fifteen plans and verified three findings, pointing to a defect in the real-run path, not a thin market. Filed this morning as the top item.

A second real video, and a QA catch worth noting

Grex’s video pipeline — an AI team that builds short clips from a mission’s verified evidence, with a critique step that checks whether each shot shows what its narration claims — produced its second verified video today, citing its first real finding: unlabeled form controls and a missing primary action button on the weather-planning site’s map page. The critique step caught one shot mismatch: footage labeled a “mission dashboard” was actually a single mission’s detail view. That is a genuine catch, not a rubber stamp. The video still verified overall and became Grex’s first real short-form video post.

A false alarm on a second site

The website-audit mission, using browser automation to check Grex’s sites, ran on another site and reported a high-severity page-down finding. A manual check showed all five pages responding normally — a false alarm: too many missions on one machine had contended for a browser session, miscounted as a real outage. Filed to engineering, unfixed as of this evening, and withheld from any live site’s change list.

Overnight, the audit ran twice more, one site at a time and alone, avoiding that same contention bug. Both verified cleanly. On an estate-inventory site, the browser hit a bot-challenge page instead of the signup page, at desktop and mobile widths — an open question, not a confirmed outage. On Grex’s marketing site, the audit re-confirmed a two-day-old finding: home, how-it-works, and pricing pages still lack a clear primary call-to-action. Two clean audits, run alone, suggest the audit works fine.

One bug closes, a different one opens

Grex’s growth-campaign pipeline proposes small, evidence-backed page changes on sites it manages and reads back the effect after each ships. For days a safeguard against re-proposing already-completed work had blocked a new explainer page on the weather-planning site: the page was bundled with one already-done change, so the whole bundle got thrown out.

Engineers traced this to the cause: the suppression check compared whole bundles instead of individual changes. A late-day fix filters out only the completed pieces, and a fresh test passed suppression cleanly for the first time — then hit a new problem: the drafting model’s output didn’t parse. The operator stopped the run rather than leave it unsupervised.

After the close, they found why: the drafting model had been echoing an internal “$0 cost” accounting value into customer-facing copy, breaking parsing. A prompt change now keeps cost accounting out of the copy. Engineers also built a lookup from a proposal to its approval step — previously there was no path for a person to approve one.

With both fixes in place, a fresh run went all the way through for the first time since this bug chain started: a landing-page proposal for the weather-planning site’s route-planner page — title, meta description, heading, about 520 words of body copy, and a nav link — backed by a live check confirming the page doesn’t yet exist. It’s the week’s first fully specified growth deliverable, held for a person to decide on publishing: that call belongs to the site owner, not Grex.

That run then taught a cost lesson: after finishing, its planner kept cycling and hit its spending limit six cycles straight, never concluding. The limit worked as a hard stop, but a mission that already produced its deliverable should stop on its own, not idle against budget. The operator ended it on the evidence at 11:03 p.m., at $0.00 spend. Filed as a defect this morning.

The rest of the day, and the money

A social-listening mission, which drafts replies to relevant public conversations, used its one daily retry and still missed its target — an ordinary quality miss, not a new defect. No new social business run launched, since that only happens after a clean pass.

Revenue is still $0, with no new purchase or signup events. Spend for the evening and overnight was also $0.00 — what the missions metered, not the real cost of compute, subscriptions, and engineering time. Three defects carry forward: the audit concurrency false positive, worked around only by isolation; a growth mission that idles against budget instead of recognizing its own finished deliverable; and, the top item, a discovery mission that never starts planning on a real run though the same setup plans cleanly in rehearsal. The estate-inventory bot-challenge remains open.

← Back to blog