Glasshouse Fund

The interesting engineering here is not the trading. It is everything built to make an unreliable component (a language model) produce an auditable record: a gateway that validates and retries, guardrails the model cannot argue past, a grounding gate that blocks unsupported claims, and a scoreboard that grades its predictions in public whether they were right or not.

One daily cycle

Twice each weekday the graph runs 22 nodes end to end. Two edges matter more than the happy path: the branch that skips execution on a no-trade day, and the fact that every node is wrapped, so a failure anywhere still produces a journalled run rather than silence.

Contextualize mark-to-market research · memory Decide bull · bear · risk → PM synthesis Ground claims checked against the facts Guardrails risk review rebalance check Execute approval gate fills · mark P&L Record journal · calls site · X · memory nothing approved → skip execution, still record the day's calls any node raises finalize_failure status recorded, partial run still journalled, exit code non-zero
Every node is wrapped by guarded_node: a raise is caught, recorded against the run, and short-circuits the rest of the graph to the finalizer. The run still writes its journal: a failed day is a visible row, not a gap.

How a proposed trade survives

The model can propose anything. What reaches the book is whatever survives a fixed sequence of deterministic checks, none of which the model can talk its way past, because none of them are written in English.

PM proposes a decision per researched name BUY / SELL / HOLD, each with a stated confidence system exits injected stop-loss 15% · take-profit 40% HOLD → skipped, never a “trade” confidence ≥ 0.60 daily turnover ≤ 20% of book sector exposure ≤ 40% order ≤ 10% of book filled at the last marked price rejected capped or rejected capped or rejected shares reduced not executed every rejection is journalled with a reason injected no-ops floor budget concentration sizing
The first four gates run in RiskManagerAgent.review; the last is enforced by the execution engine. Note it caps the order, not the resulting position: a holding already near the limit can still be topped up, which is a known gap rather than a design choice. Rebalance trades re-enter at the top of this same sequence; there is no bypass.

System architecture

One rule shapes the layout: every model call goes through a single gateway, and nothing that decides money is written in English.

TRIGGERS cron ×2 weekdays health watchdog 22:15 ORCHESTRATION LangGraph: 22 guarded nodes resume by run_id · progress in SQLite AGENTS analysts · portfolio manager · risk · rebalancer letter writer · tweet writer · grounding judge GATEWAY one LLM gateway strong (terra) → decisions, judges · cheap (luna) → analysts, tweets schema validation + repair · backoff · 60s timeout PROVIDERS OpenAI gpt-5.6-terra · gpt-5.6-luna yfinance · news · SEC fallback provider slot unset every model call Qdrant memory theses · lessons · trades · macro read + write stores decisions · trades · predictions letters · history · run log append-only jsonl / csv / json append static site prerendered, on Pages X / @GlassHouseFund 3 posts per run MCP server read-only, 6 tools tweet writer drafts each post reads evals in CI golden decision set · ablations (memory / debate) · chunking · grounding temperature 0, no network, runs on every pull request
The gateway is the highest-leverage piece: routing, retries, structured-output validation, timeouts, tracing and cost accounting are paid for once rather than per agent. Everything downstream of it is deterministic.

Failure is the normal case

A daily job that depends on a scheduler, three external APIs and a language model will fail. The design assumption is not that it won't; it is that a failure must never look like a success.

FailureWhat catches it
A node raises mid-runguarded_node records it on the run, skips the remaining nodes, and still journals the day. The process exits non-zero so CI goes red.
A model call stallsA 60-second client timeout, then the gateway's exponential backoff. Added after one stalled connection froze a batch for ten minutes.
A model returns malformed JSONThe gateway re-prompts once with the validation error, then fails loudly rather than passing a half-parsed decision downstream.
The model states something it cannot supportA grounding judge compares every claim against the facts the writer was given. Ungrounded decisions and letters are blocked before publication, not after.
A process dies halfwayRe-entering reuses the same run_id; the stores dedupe on it, so replaying a partial run does not double-write or double-tweet.
The run never happens at allAn independent watchdog. A run cancelled before its first step writes no row anywhere: the audit trail cannot record its own absence, so only an outside observer can notice the gap.

What is actually measured

Claiming an AI system works is cheap. These are the numbers it is held to, all of them regenerated every run.

MeasureMethod
Prediction accuracyEvery researched name gets a directional call against the S&P at 5- and 30-day horizons, recorded before the outcome is known and scored automatically when the window closes. Decoupled from trading, so the sample is not limited to names the fund bought.
CalibrationStated confidence bucketed against realised hit rate, plus a Brier score. This is the measure the project exists to publish: whether the model's confidence means anything.
Component valueAn ablation harness scores the full system against no-memory and no-debate variants on a fixed eval set with a fixed judge: evidence that the AI parts earn their place, rather than an assertion.
BaselinesThe fund is compared against buy-and-hold SPY, 100% QQQ, and the mean of 500 random equal-weight portfolios drawn from its own watchlist.
CostEvery model call is logged with tokens, latency and dollar cost. A measured finding: the flagship models cost roughly 11× more for no reliable gain in decision quality on this eval set, so the strong tier is deliberately not the biggest model available.

Key engineering decisions

The calls that shaped the project, and the reasoning behind them.

DecisionWhy
One LLM gateway as the single choke point The highest-leverage refactor in the repo. Routing, retries, structured-output validation, tracing and cost tracking become one-time costs instead of per-agent ones. Every agent flows through it, and the provider seam was built on it.
Deterministic guardrails outside the model The portfolio manager is allowed to be creative; the risk layer is boring on purpose. Position sizing, turnover, sector limits and stop-loss/take-profit are enforced in code the model cannot argue past.
Append-only files as the system of record Decisions, trades, predictions, letters and run history are JSONL, CSV and JSON: diffable, replayable, and readable without a client. SQLite is used for exactly one thing: the run-progress table that makes resume work. Raw state is committed alongside the site, so the published numbers and the data behind them move together.
LangGraph as the only runner The legacy linear orchestrator was deleted once the graph reached parity. Two orchestrators is debt, not safety.
Idempotency-based resume, not native checkpointing Graph state carries live, non-serializable handles. Persisting progress and reusing the run_id leans on already-idempotent stores for duplicate-free re-execution: correct, without a large and risky state refactor.
Prerendered pages, not a client-rendered app The journal used to be one URL that fetched JSON, so a crawler saw almost nothing. Every trading day, every ticker and every letter is now a static page with real text at a permanent URL.
Grounding gate before anything publishes Daily decisions and weekly letters are both checked against the facts they were given. It has real limits: it verifies that a number came from the fact base, not that the fact was labelled correctly: a mislabelled field passed the judge and reached four published letters before it was caught.
Publish the misses too Rejected trades, blocked letters, failed runs and wrong predictions are all part of the public record. A scoreboard that only shows wins is marketing; the point of this project is that it is checkable.