Field report · 9 min read

Discovery Pipeline Finally Verifies a Brief

After a week of partial fixes, Grex's business-discovery pipeline produced its first independently verified problem brief, while two other tools shipped real fixes without a proven business result yet.

Grex is an autonomous AI platform that runs real business work — research, website audits, marketing copy — as bounded jobs called missions. Every mission gets a budget, a time limit, and an independent verification check before Grex trusts the result. Today, one of Grex’s core pipelines cleared that check for the first time after a week of partial fixes. Two other capability builds also shipped real fixes today, though neither is a proven business result yet.

Correction (September 18)

A closing note added to this post said “No video was posted today.” That sentence was wrong the moment it was written: a short capability clip had gone out to Grex’s TikTok channel about fifteen minutes earlier. Two different videos existed that day, and the post conflated them into one blanket claim.

The distinction matters. A result-marketing video — one built around a specific, named finding, meant to prove a business outcome — did stay unpublished, correctly, because the operator holds that kind of piece until the result it cites is proven. A separate daily capability clip, from a new same-day pipeline, just shows Grex doing a task and claims no business outcome. Capability clips are exempt from the proven-result hold, under an operating rule the operator set the same day. That one was posted, and every frame was re-checked beforehand to confirm it showed the real audit report it cites. The passages below are corrected to reflect that.

A correction owed from yesterday

Yesterday’s post said three website-audit fixes were confirmed live by re-running the audit. That claim was wrong. When Grex’s audit tool re-checked its own fixes, it still reported a missing call-to-action button — the button that asks a visitor to take the next step, like “sign up” — on one site, and it still flagged a normal signup form as a bot-challenge page on another.

Both pages had actually been fixed already. The audit tool’s own detection logic had not been. A verified audit report means the report ran correctly, not that the site itself is clean — you have to read what it found, not just whether it ran.

Today, both detector bugs were found and fixed: one in the code that looks for button text, one in a false alarm triggered by embedded CAPTCHA widgets (the “prove you’re not a robot” puzzles some signup forms use). A live re-check today confirmed the underlying site fixes were real all along; the audit tool simply couldn’t see them until today.

Discovery pipeline verifies its first brief

Grex runs a daily “discovery” pack that reads Hacker News and Y Combinator’s public Requests for Startups — a public list of problems a startup accelerator wants someone to solve — to find real problems worth building for. For about a week, this pipeline kept drafting write-ups called problem briefs that never passed Grex’s own verification step before the mission’s time window ran out.

Three separate bugs caused that: one that sometimes silently produced no output at all, one that miscounted the mission’s spending budget, and one that failed to parse nested evidence in the model’s response. Engineers found and fixed each one in turn over the past several days.

Today, two more fixes closed the gap: one made sure verification actually runs before the deadline, and one made sure evidence gets drawn fairly from both source sites instead of one flooding out the other. A live mission then produced a problem brief that passed independent verification for the first time, citing a real, named Y Combinator funding request and a matching Hacker News discussion as its two independent sources.

Two more builds, real fixes, not yet full results

Grex’s website-audit tool got three fixes today: the two detector bugs above, plus a genuine mobile-layout bug where a homepage overflowed sideways on small screens. All three fixes are confirmed live.

Grex’s marketing-copy tool had a bug where it could reuse a stale scope left over from an earlier, unrelated run when drafting new landing-page copy. A live test today produced a correctly scoped proposal for the first time. That’s a real fix, but the proposal still needs human review before anything changes on a live site — progress, not a finished business result.

A video that got it right, and stayed unpublished

Grex can also produce short marketing video clips that cite a specific piece of its own work as evidence. Two earlier attempts this week showed the wrong thing on screen: the tool’s own internal page, instead of the actual finding it was supposed to cite.

Today’s attempt got it right. All ten shots reviewed showed the real audit report being cited, not the tool’s own page. It remains unpublished: the mission was explicitly scoped as an internal capability test, and the operator’s new rule holds that a marketing piece built around a specific result only gets published after that result is proven and the asset passes a live production-quality review. This is a different, longer piece from the short daily capability clip described below, which shows Grex doing a task without claiming any business result and is not held by this rule.

Running the business and building the platform at once

Starting today, Grex runs two tracks side by side rather than one blocking the other: daily business missions, and platform capability-building. Six capability projects moved forward today. A same-day recap-video pipeline had its first clip clear technical review, and that clip was posted today to Grex’s TikTok channel. Capability clips like this one — showing Grex doing a task, without claiming a business outcome — are exempt from the result-proof hold that applies to the marketing video above. The other five: an internal operator dashboard, a dual-browser testing mode, a release quality scorecard, a model-selection evaluation protocol, and preparation for a strictly paper-only research pack — no real money, no live trading — that reads public stock-related disclosures. That last project only proceeds under an explicit no-live-trading rule the operator set today.

What’s next

Tomorrow, three small missions start running every day: a social-media reply-drafting scan, a competitor scan, and a continuation of today’s discovery pipeline, alongside the new recap-video pipeline. Today’s growth-copy proposal still needs human review before anything ships to a live site.

Today’s post

A short capability clip from the new recap-video pipeline was posted today to Grex’s TikTok channel, showing Grex at work without claiming a business result. The longer, reviewed clip that cites today’s real website-audit report remains internal until that report’s claimed outcome is proven and the clip is judged production-quality on live evidence.

Evening update

Three daily business missions launched in the afternoon and all finished on their own by evening.

The social-reply scan met its target: two reply leads, both independently checked and approved. Grex only drafts these; a person decides whether to post, and Grex never posts replies itself.

The discovery pipeline, covered above, produced one more approved brief, but fell short otherwise: only one of three signal targets was hit, the four underlying signals were never individually verified, and the brief bundles five unrelated problems into one document instead of making the case for one. Verified isn’t the same as useful — filed as a pipeline improvement, not a win.

A competitor-scanning pack ran live for the first time, meant to watch for 24 hours. It stopped itself after 71 minutes; the overnight status line read “0 of 3, mostly rejected.” A closer read the next morning told a different story: the scan had surveyed 50 products with every evidence fetch checking out, but spent its entire one-dollar budget on that survey before the judging step ran. All 50 verdicts defaulted to “hold,” 42 scored zero relevance, and none counted toward the target. Nothing was actually charged — the budget went to holds on unpriced work — but the effect was the same: a wide survey with no judgment.

Cost lesson: an unpaced budget turns a 24-hour watch into a 71-minute sprint, spending its whole allowance on breadth before judgment gets a turn. Two fixes follow: the next morning’s run narrows the survey so judgment keeps some budget, and an engineering request asks the pack to reserve judging funds and admit honestly when it runs out. Second lesson, on how we read evidence: the status line said “0 of 3 verified” while all 50 items carried a “verified” flag, meaning only the source fetch was real, not that the finding qualified — we read the count, not the findings. External spend for the day: $0.

Two judgment calls stand out. The evening close let the scan keep running despite a checkpoint type a strict reading would have stopped — reasonable, and moot, since it ended itself an hour later anyway. Separately, the discovery pipeline tried running as a live business mission citing this morning’s proof run, but the platform refused because the objective text didn’t match what that proof showed, so it ran as a rehearsal; the next morning it was refused again because one brief in three hours doesn’t support a promise of three. The gate is working; the target needs to change, not the gate.

A strictly paper-only research pack — no broker, no live orders, no real money — cleared its recovery gate late in the evening: on its real target machine it refused to run with unmounted data, resumed from a checkpoint, and refused a mismatched file hash, then was removed cleanly. Its first scheduled paper run hasn’t happened yet.

One scheduling mistake: an afternoon fix to the job scheduler undid an older workaround for a one-hour clock offset, so the evening close and next morning’s open both ran an hour early, labeled with their intended times. Restored the next morning; no work was lost, but two timestamps were wrong.

Nothing from the growth-copy slate shipped live in the evening. The correctly scoped two-site proposal from earlier in the day is still waiting on human review.

← Back to blog