Field report · 6 min read

Nine attempts to verify one paper stock decision

On day two of four daily AI jobs, Grex verified one paper decision and 14 competitor briefs, yet met none of its targets.

Grex is an AI team platform. It takes bounded assignments called missions, does business work, and shows its evidence. This week it runs four daily jobs. On day two, Grex verified its first paper stock decision but met none of its targets.

“Verified” means an independent check confirmed an output against its source evidence. A mission has a “planning step” that decides what to do and a “judging step” that reviews or drafts the result. “Paper trading” means a simulated ledger with no orders, no broker, and no real money.

The four jobs at a glance

Job Result by the end of the day
Social watch 0 of 2 verified, in two 3-hour runs; benched
Competitor scan 14 briefs verified, none counted; run failed
Ideation 0 of 3 signals verified; business run refused
Stock research, paper only First verified paper decision on attempt 9; seven-day series started at 18:19

Social watch missed a second day

Social watch finds Bluesky posts worth a reply and drafts the reply for a person. Both runs ended with 0 verified against a target of 2. The planning step made about 104 model calls across the two runs. The judging step made none, so nothing reached a check. Engineers benched the job until they fix that class of defect.

The competitor scan verified briefs that did not count

The scan tracks what moved in the agentic-AI space. Its 24-hour run ended at 17:00 when it missed its second checkpoint of 3 verified briefs. It produced 16 briefs, and an independent check verified 14 and could not verify 2.

Of the 14, 13 received a “hold” verdict as adjacent products. One received “enter” as a direct competitor: an assistant that proposes actions and drafts messages for approval, scoring 75 of 100. None counted toward the target, so the run is recorded as failed.

An operator read all 16 briefs that evening, and the failure is deserved. The target counts briefs that pass a second review called ratification, and none did. More important, the briefs were mostly noise: a machine-learning experiment tracker and six unrelated code repositories. The one “enter” verdict is a sales-outreach tool where a person approves every message. It does not compete with Grex. The assignment named two real competitors, and the scan never attempted either. “Verified” here meant the quotes matched their sources. It did not mean the product was relevant. The request to engineers changed from “make these count” to “attempt the named targets first.”

Ideation was blocked by a freshness rule

A 3-hour rehearsal made one problem brief, but it could not be verified. A business run was then refused because the previous day’s rehearsal was more than 24 hours old, by minutes. A daily cadence cannot meet that rule.

Stock research finally got a verified decision

Attempt 8 again produced a paper “abstain” decision, because its data folder was not mounted. It stayed unverified. Engineers deployed a fix, and attempt 9 at 13:08 produced the first verified paper decision, also “abstain”, at $0.

That met the exit condition, and engineers released the go-ahead. The first seven-day launch was refused, because an operator had stopped attempt 9 early and a stopped run cannot certify the next one. So a four-hour rehearsal started at 14:16 and was left alone until it ended by itself at 18:16 with one verified decision. The seven-day series launched at 18:19, citing that result.

Five hours later the series stopped planning. Its budget of 40 model calls had been copied from the four-hour rehearsal. This job replans every 15 minutes, so a seven-day run needs hundreds of plans. The planner uses fixed rules and had made no model calls at all, yet each plan was counted as one. Engineers changed the counter to count a call only when tokens or spend exist. Planning resumed at midnight without a relaunch. The first real trading-day slot is Monday at 07:00 Central, and the 12-month paper clock starts with it.

A search-demand study did not finish

Operators asked whether search terms for three small consumer sites are winnable. Keyword volumes were measured, but page-ranking snapshots were refused and a shared model quota ran out. Two retries returned empty plans. Five attempts ran by 21:15, and none produced a verified result.

The last attempt shows a cost trap. It had a $1 cap and ended one hour into a three-hour window with its budget exhausted and $0.00 spent. Each scheduled run reserves money before it works. The runs had no plan, so they did nothing, quickly, and the reservations alone passed the cap.

What the counts prove

At 18:00 the report cards showed 0 judging calls on every mission, and engineers filed that as one class of bug. Overnight that number became less certain. On a benchmark run, the same counter recorded 42 of 1,134 model calls, because it credits calls by worker name rather than by role. Some “zero judging” readings may be a counting error. The social watch and competitor scan failures stand on other evidence: off-topic candidates and unattempted targets. Engineers, not Grex, wrote every fix described here.

Two budgets that counted the wrong thing

Both budget failures have one cause. The dollar cap and the call budget meter scheduled runs, not useful work. A $1 cap ended a mission that spent nothing. A 40-call budget stopped a planner that called no model. The rule we took from it: size every budget from the new window and the job’s cadence, and never copy one from a shorter run.

A ruling on which models judge

Several missions share one subscription quota for their judging model. A benchmark run drained it at 13:02, and judging calls failed for an hour. Around midnight the operator ruled that judging moves to a cheaper model mix, with an independent model family as the checker. An overnight benchmark on that mix is running. It has no result yet, and there is still no comparison run against a plain loop on the same model.

No mission produced a verified source, so no capability clip was posted. Measured external spend was $0.00. That is not zero cost, since subscriptions, hardware, and engineering time are uncounted. Some missions held $1 to $3 authorizations, and holds are not spend.

The counts show real progress in one place: a verified paper decision and a seven-day series that is running. The 14 verified briefs are not progress, because they were about the wrong products. They do not show a met target, and a correct stop does not make a missed mission a success.

← Back to blog