How the fund is built
An LLM runs a simulated $1M book in public. A 22-node LangGraph cycle sits on one hardened LLM gateway; deterministic guardrails sit outside the model; append-only files are the system of record; and every directional call the fund makes is scored against what actually happened.
The interesting engineering here is not the trading. It is everything built to make an unreliable component (a language model) produce an auditable record: a gateway that validates and retries, guardrails the model cannot argue past, a grounding gate that blocks unsupported claims, and a scoreboard that grades its predictions in public whether they were right or not.
One daily cycle
Twice each weekday the graph runs 22 nodes end to end. Two edges matter more than the happy path: the branch that skips execution on a no-trade day, and the fact that every node is wrapped, so a failure anywhere still produces a journalled run rather than silence.
guarded_node: a raise is caught, recorded against the run, and short-circuits the rest of the graph to the finalizer. The run still writes its journal: a failed day is a visible row, not a gap.How a proposed trade survives
The model can propose anything. What reaches the book is whatever survives a fixed sequence of deterministic checks, none of which the model can talk its way past, because none of them are written in English.
RiskManagerAgent.review; the last is enforced by the execution engine. Note it caps the order, not the resulting position: a holding already near the limit can still be topped up, which is a known gap rather than a design choice. Rebalance trades re-enter at the top of this same sequence; there is no bypass.System architecture
One rule shapes the layout: every model call goes through a single gateway, and nothing that decides money is written in English.
Failure is the normal case
A daily job that depends on a scheduler, three external APIs and a language model will fail. The design assumption is not that it won't; it is that a failure must never look like a success.
| Failure | What catches it |
|---|---|
| A node raises mid-run | guarded_node records it on the run, skips the remaining nodes, and still journals the day. The process exits non-zero so CI goes red. |
| A model call stalls | A 60-second client timeout, then the gateway's exponential backoff. Added after one stalled connection froze a batch for ten minutes. |
| A model returns malformed JSON | The gateway re-prompts once with the validation error, then fails loudly rather than passing a half-parsed decision downstream. |
| The model states something it cannot support | A grounding judge compares every claim against the facts the writer was given. Ungrounded decisions and letters are blocked before publication, not after. |
| A process dies halfway | Re-entering reuses the same run_id; the stores dedupe on it, so replaying a partial run does not double-write or double-tweet. |
| The run never happens at all | An independent watchdog. A run cancelled before its first step writes no row anywhere: the audit trail cannot record its own absence, so only an outside observer can notice the gap. |
What is actually measured
Claiming an AI system works is cheap. These are the numbers it is held to, all of them regenerated every run.
| Measure | Method |
|---|---|
| Prediction accuracy | Every researched name gets a directional call against the S&P at 5- and 30-day horizons, recorded before the outcome is known and scored automatically when the window closes. Decoupled from trading, so the sample is not limited to names the fund bought. |
| Calibration | Stated confidence bucketed against realised hit rate, plus a Brier score. This is the measure the project exists to publish: whether the model's confidence means anything. |
| Component value | An ablation harness scores the full system against no-memory and no-debate variants on a fixed eval set with a fixed judge: evidence that the AI parts earn their place, rather than an assertion. |
| Baselines | The fund is compared against buy-and-hold SPY, 100% QQQ, and the mean of 500 random equal-weight portfolios drawn from its own watchlist. |
| Cost | Every model call is logged with tokens, latency and dollar cost. A measured finding: the flagship models cost roughly 11× more for no reliable gain in decision quality on this eval set, so the strong tier is deliberately not the biggest model available. |
Key engineering decisions
The calls that shaped the project, and the reasoning behind them.
| Decision | Why |
|---|---|
| One LLM gateway as the single choke point | The highest-leverage refactor in the repo. Routing, retries, structured-output validation, tracing and cost tracking become one-time costs instead of per-agent ones. Every agent flows through it, and the provider seam was built on it. |
| Deterministic guardrails outside the model | The portfolio manager is allowed to be creative; the risk layer is boring on purpose. Position sizing, turnover, sector limits and stop-loss/take-profit are enforced in code the model cannot argue past. |
| Append-only files as the system of record | Decisions, trades, predictions, letters and run history are JSONL, CSV and JSON: diffable, replayable, and readable without a client. SQLite is used for exactly one thing: the run-progress table that makes resume work. Raw state is committed alongside the site, so the published numbers and the data behind them move together. |
| LangGraph as the only runner | The legacy linear orchestrator was deleted once the graph reached parity. Two orchestrators is debt, not safety. |
| Idempotency-based resume, not native checkpointing | Graph state carries live, non-serializable handles. Persisting progress and reusing the run_id leans on already-idempotent stores for duplicate-free re-execution: correct, without a large and risky state refactor. |
| Prerendered pages, not a client-rendered app | The journal used to be one URL that fetched JSON, so a crawler saw almost nothing. Every trading day, every ticker and every letter is now a static page with real text at a permanent URL. |
| Grounding gate before anything publishes | Daily decisions and weekly letters are both checked against the facts they were given. It has real limits: it verifies that a number came from the fact base, not that the fact was labelled correctly: a mislabelled field passed the judge and reached four published letters before it was caught. |
| Publish the misses too | Rejected trades, blocked letters, failed runs and wrong predictions are all part of the public record. A scoreboard that only shows wins is marketing; the point of this project is that it is checkable. |