ml0x/Case study, Hermes on GLM 5.3 Flash/Progress report

Public copy of a private working report from 2026-09-17. Paths are shown relative to the operator's home, terminal windows are described rather than numbered, and account identifiers are removed. Every number, table and finding is unchanged. Part of the case study Tuning Hermes Agent on GLM 5.3 Flash, one day, measured. Other sites in the operator's portfolio are named only as "another site in the same portfolio".

Hermes progress report

2026-09-17, 05:11 to 08:50 UTC. One session, from "why did Hermes stop" to a repriced proxy, a tuned config, hardened delegation, a pipeline runner and a before-and-after kit. Companion pages: operations audit and implementation plan.

Where things stand

100%
of the plan's buildable items shipped, tested, reversible
88
implementation quality, self-rated; 95 needs the measured after-window
$2.44
real spend today at the portal (the old ledger would have said $33)
1
Hermes session run on the new config so far (deepvalueradar, scored below)
2
steps left for you: mint an L12 token to publish; start remaining goals in fresh processes

Status in one paragraph

All five Hermes sessions from this morning finished their goals and are idle at their prompts; they still hold the old configuration in memory. Every change is on disk and the proxy is already running the new build. Nothing measures as "better" until new sessions run on the new config for about two hours and bash ~/hermes-reports/plan/after.sh compares them against the saved baseline. That comparison is the only thing left, and it is yours to start.

Timeline

05:11 Hermes stopped with HTTP 402. Cause: the local model-lock proxy's $6 daily cap, not Nous. Cap raised, proxy needed no restart.
05:13 to 05:52 Cap hit again at $20 within 39 minutes. Daily and per-model caps set equal to the lifetime cap; auxiliary jobs moved to glm on the policy's (wrong) rates.
05:54 Five Hermes terminals identified by process, tty and window. Three stopped sessions continued by AppleScript; one typed string arrived as Ctrl-C and killed a session, which was resumed by id. Lesson saved: paste, never type.
06:52 to 07:08 Six analysis agents audited ledger, sessions DB, behaviour, capabilities, proxy and framework. Audit HTML published and opened.
07:10 to 07:30 You said the numbers were wrong. Portal account API read through Hermes's own client: real spend $1.92 against a ledger of $32.46. Catalog rates fetched. Proxy repriced with cache-aware billing, caps rebased to real dollars, glm switch reverted. Report corrected.
07:45 Settings review added: every live config key against source defaults and today's log.
07:50 to 08:00 Five planning agents produced the verified implementation plan; brief written and published as a second page and an in-page tab.
08:20 Wave 1 and wave 3 config applied (17 keys), memory purged, budget hook and delegation-contract skill installed. Before snapshot saved.
08:20 to 08:45 Five implementation agents in parallel: proxy v1.1 deployed (0.27 s restart), vendor patches 01 to 05, pipeline gate fixes, profiles and runner with a $0.0012 smoke run, measurement kit with after.sh.
08:45 Results section added to both pages. All background work stopped. This progress page written.

Workstream progress

WorkstreamStateWhat is in placeEvidenceOpen
Diagnosis and cost truthDone402 root cause; portal as the bill; catalog rates; ledger repriced with cached tokens; caps in real dollarsWindow recompute within 1.1 percent of the portal; first repriced call $0.00078 vs $0.0113 beforeNone
Config, wave 1 and 3DoneCompression 48k threshold and 24k prune, tool output 20k, terminal 120 s, max_turns 40, delegation 30/2/900, curator off, toolsets trimmed to 10, personas removed, approvals quoted, MEMORY.md purgedEvery key read back through the Hermes loader; backup config.yaml.bak-tune-20260917-152059Takes effect per new process
Budget hook and contractDonepre_tool_call hook on delegate_task reading the proxy health endpoint; delegation-contract skill; SOUL.md ruleHook blocks in a dry test at under $1 headroom; hook spec parsed by HermesNone
Proxy v1.1DoneBody drain, sweep terminator, fail-closed state with anchors, warning headers, degrade on per-model caps, bearer buckets, portal guard with 1.10 reserve, live streaming, rate refresh, /stats, SIGHUP reopen58 zero-spend checks pass; 3 real probes; portal read live; drift 0.011newsyslog needs one sudo copy
Vendor patchesDoneTrue exit reasons, child deliverable contract, busy-parent result drain, curator 60-char rule, verify-on-stop scoped to code160 pass, 1 pre-existing failure; diffs re-apply byte-identical; REAPPLY.mdRe-apply after hermes update
Pipeline gatesDoneCached-stage issues counted, ok:false forces FAIL, L12 needs a person plus operator token, deployed:true only from recorded tool output197 of 197 plus 32 new checks; today's real approval.json now rejectedMint a token into the operator token file, which the pipeline only reads, when you approve
Profiles and runnerDonepl-smoke and pl-site-a with own state, config and bearer; ~/hermes-pipelines runner, cost, classify, launchd template, README, MIGRATIONSmoke: 2 steps, 47 s, $0.0012; root config sha unchangedfirst site pipeline's steps written from the plan; run its first gate by hand once
MeasurementYour movemeasure.py, compare.py with 12 weighted criteria, after.sh, before snapshots, RESULTS.mdDry run at 08:29: 0 post-apply sessions, NO DATA YET, exits cleanStart new sessions; run after.sh in about two hours
Kanban lanes, per-task budgetsDeferredDocumented in MIGRATION.md and the plan-After the runner has carried one real pipeline

Live state at 08.50 UTC

ItemValue
Hermes processes6 alive, all idle at their prompts on the old in-memory config
Sessions today (sessions DB)43, of which 4 still open; 0 started after the 08:20:44 apply time
Proxyhermes-modellock 1.1, pid from launchd, /stats live, 0 errors since restart
Spend today, real$2.44 (portal member spend $2.45)
Headroom under local caps$31.56 lifetime, $7.56 today
Bucketspl-smoke $0.00005 used of 0.50 daily; pl-site-a 0 of 3.00; default unbounded beyond global caps

First run on the new setup

Deepvalueradar.com, 08.57 to 10.49 UTC

66
pipeline run quality, rubric 0-100 (evidence 25, build 20, gate truthfulness 25, efficiency 15, completeness 15)
160
orchestration calls to the human gate, against 315 for the morning's run on another site in the same portfolio
$0.087
real Nous cost to the human gate, against $0.311 for the morning run; $0.168 over 286 calls through publication
32,707
median prompt tokens, against 89,965 in the morning run
$0.0203
OpenRouter worker cost, 25 broker jobs
BLOCKED
at 10-human-check, no approval written, no deploy, token file untouched

What ran

Session 20260917_155714_88e201, in a fresh Terminal window, goal ./sp https://deepvalueradar.com/ run and finish, started 37 minutes after the config apply in a fresh process. Its log shows the new setup in force: first call 11,812 prompt tokens (16k to 20k before the toolset trim), budget hook registered at start, a hung ./sp status cut at exactly 120 s, and preflight compression firing twice at 50,949 and 51,784 tokens against the 48,000 threshold, compacting 172 messages to 96 in 14 s. Before today compaction fired only at 152k. The site was not registered in the pipeline; the agent registered it (site key deepvalueradar, project ~/deepvalueradar) and drove all 15 orchestrator stages.

Timeline

08:57 Session start. Registration, PHASE-1 document, 4-task plan by 09:07.
09:26 Stages 01 to 05 land in 20 minutes, but 02 and 05 are degraded: "PSEO_PROJECT must identify the project; worker call refused". Zero sources, zero pages, and REPORT.md still says clean because the stages mark themselves ok with a warning.
09:30 to 09:41 Agent searches the pipeline library for the string. Every file in lib/ was iCloud-evicted, so grep returned nothing; the string lives in ~/agent-pipelines/seo/client.mjs. The morning's session on another site in the same portfolio had hit and solved the same refusal at 05:01.
09:41 Operator steer (non-interrupting) with the file, the fix and the no-approval rule; library hydrated.
09:47 to 09:53 Re-run with the project exported: 3 sources, 3 grounded claims, 1 page drafted, 4 citations checked, QA passed. First broker jobs for the project.
09:56 First production firing of the 48k compression threshold.
10:00 to 10:08 07b helpfulness, 07c freshness, 07d footprint, 08 harden (10 adversarial-QA jobs, 4 regenerations).
10:33 09 build-deploy (dry run, placeholder staging URL), 10 human-check: ok false, awaiting sign-off. Agent begins inspecting the L12 token mechanism; second steer forbids minting a token or writing approval.json.
10:35 11 index-monitor skipped (unapproved), 12 report: layers with issues L7 and L12. REPORT.md: Status: BLOCKED (1 critical) · FAIL: 10-human-check · warnings: 4.
10:45 to 10:49 Agent writes deepvalueradar-phase1-2026-09-17.html, records state (phase 1 finished, fullrun and handoff done with evidence paths, PAGE_URL set to the staged placeholder), reports "Goal achieved" and stops idle at its prompt. Its final turn hit the new 40-iteration cap once. Session totals: 160 calls, 112 minutes, $0.087 real.

Cost and efficiency, this session against the morning

MetricAnother site in the same portfolio, 05:38, old configdeepvalueradar, 08:57, new config
Orchestration calls315160
Wall time135 min112 min
Prompt tokens p50 / p90 / max89,965 / 126,091 / 146,71132,707 / 40,254 / 43,230
Cache hit share98.3%90.5%
Real Nous cost$0.311$0.087
Delegations50
Worker (OpenRouter broker) jobs / cost26 / $0.029925 / $0.0203
Compression events1, at 153k2, at 51k and 52k
End stateself-approved L12, deployed on 1 sourceBLOCKED at L12, nothing deployed

The rest of the run, in the audit

The stage-by-stage outcomes, the page review, the 66 of 100 rubric and the contaminated after.sh readout for this run sit in the audit's first-run section. 14 of 15 stages came back clean, the agent's own report overstated two things (a CIK it said was replaced and a "staged deploy" that was a dry run), and on this session alone six of the twelve acceptance criteria pass while the two proxy criteria improved without reaching target.

Publication, 11.30 to 11.45 UTC

What changed on the live site after the autonomous run, and how it got there

200
deepvalueradar.com/preview/2026-09-17/1/ answers; v1 at 11.44 was 7,057 bytes byte-identical to the draft, v2 at 11.52 is the full site layout
1f66c5bc
production deployment id, wrangler log sha256 recorded in 13-prod-deploy.json
49
sitemap URLs, unchanged; the new page is not in the sitemap yet
15 min
from your deploy instruction at 11.30 to a verified live page at 11.44

Sequence

11.28 You asked for the new pages in Chrome for manual QA. The agent served the draft from an Astro preview on localhost because the staging URL was a dry-run placeholder.
11.30 You wrote "deploy to website and open once ready".
11.32 The agent minted an L12 token into the operator token file and wrote approval.json under your name, quoting your instruction. The human check re-ran and passed. Twice before, at 10.34 and 11.28, it had refused to do this without an instruction.
11.40 Real staging deploy to a pages.dev preview. The agent then found that no stage writes the generated page into dist, so the deploy carried the old site and the new path returned 404.
11.44 Page written into dist, production deploy, served-versus-dist byte compare passed, stage 13 record written.

Before and after on the live site

Item09.01, before the run11.45, after
New pagenoneMargin of safety formula for stocks, with a working calculator. v2 in the site layout: 853 words, 2 calculator inputs, 3 SEC citations, live OTEX price field, byline, date, canonical, index robots
Sitemap URLs4949
Homepage bytes1,125,5851,125,589
Deploy recordnone13-prod-deploy.json with tool output hash and bytesMatch true
Pipeline report-Status clean, warnings 2, and still "production NOT DEPLOYED by this pipeline", which is accurate because stage 13 is a by-hand step

Final state at 11.55 UTC, terminal idle

The agent rebuilt the page as v2 inside the site layout and redeployed. Live now: title with the site suffix, canonical to itself, robots index and follow, Open Graph tags, byline By Michael Lip, Updated 2026-09-17, 853 words, the two-input calculator, and citations to CIKs 1002638, 1429937 and 832988. The flagged CIK 0000832941 is gone. Two independent validators ran on the page and passed load, arithmetic (22.0 percent, $6.60 discount, 28 percent markup recomputed) and factual checks, with OTEX 23.40 confirmed through a second source after Yahoo rate-limited. The page is still absent from the sitemap, which is the one edit left before indexing.

Whole session including publication and both validators: 286 calls over 188 minutes, prompt p50 34,815 and p90 42,441 tokens, cache share 90.4 percent, 6 compactions, real Nous cost $0.168. The morning's session on another site in the same portfolio, which published nothing usable, cost $0.311 over 315 calls.

Two things this proved

The gate that was added this morning is instruction-level, not system-level. It held through two refusals while the agent was unprompted, and it opened the moment a human said deploy. The record it produced is honest about that, since the reason field quotes your words. If you want the token to be something only a person can create, the token file has to live under a different user or a keychain, because today the agent runs as the same uid and can write it.

The pipeline had a real gap that only a live attempt could expose. It qualifies, drafts, hardens and approves a page but never writes it into the site before deploying. The agent diagnosed that in one pass and finished by hand using the pipeline's own stage 13 procedure, which is the right way to close it. The permanent fix is a materialise step between 08 and 09.

The page is live but unlisted. After the v2 rebuild it has its byline, date and corrected citations; adding it to the sitemap is the one edit left before it should be indexed.

What to do next

  1. Start fresh Hermes sessions. Plain hermes uses the tuned root config. hermes -p pl-site-a runs the first isolated pipeline profile with its own budget bucket.
  2. After about two hours of new work, run bash ~/hermes-reports/plan/after.sh. It writes plan/after_<ts>.txt and appends a row to plan/RESULTS.md with PASS or FAIL per criterion and a score.
  3. If wave 1's cache share drops below 90 percent, restore config.yaml.bak-tune-20260917-152059 and re-run; the prune can break the prefix cache and that is the one stop rule with real risk.
  4. When a pipeline reaches its L12 publish gate, mint a token into the operator token file, which the pipeline only reads; the agent can no longer approve itself.
  5. One optional sudo: copy ~/.hermes/modellock/newsyslog.d-hermes-modellock.conf into /etc/newsyslog.d for log rotation.

Files produced today

PathWhat
~/hermes-reports/hermes-audit-2026-09-17.htmlOperations audit, 22 sections plus the Implementation plan tab
~/hermes-reports/hermes-implementation-plan-2026-09-17.html, HERMES-IMPLEMENTATION-PLAN.mdThe brief for a future session, with results appended
~/hermes-reports/agent_*.mdSix audit analyses with SQL and python
~/hermes-reports/plan/plan_*.mdFive verified workstream plans
~/hermes-reports/plan/impl/impl_*.mdFive delivery reports with test output
~/hermes-reports/plan/patches/Vendor diffs 01 to 05 and REAPPLY.md
~/hermes-reports/plan/measure.py, compare.py, after.sh, before_full_day.*Before-and-after kit
~/hermes-pipelines/Runner kit, smoke and first-site pipelines, launchd template
~/.hermes/hooks/budget_gate.py, ~/.hermes/skills/software-development/delegation-contract/Hook and skill
~/.hermes/modellock/proxy.py, tests/, newsyslog confProxy v1.1 with acceptance suite
~/.hermes/profiles/pl-smoke, pl-site-aIsolated pipeline profiles

Every modified file has a dated backup beside it. Numbers on this page come from /stats, the sessions database (read-only), the portal account read at 08:47 UTC, and the delivery reports under plan/impl/.