Building a 24/7 HFT System With Jev: The Decision Engine Is Not the Trading System

According to public materials, TypeSafe AI came out of about two years of stealth on September 15 and announced roughly $40 million in funding led by DCVC; the new model released alongside it is Jev. Reportedly, founder Diogo Almeida worked on ChatGPT and InstructGPT. For people building trading systems, brand narrative matters less than the slot it is trying to fill: not a chat model that writes long prose, but a fast decision engine that takes state as input.

According to public docs, you send market state (for example an order-book snapshot) together with typed questions; Jev returns typed answers with probabilities and confidence. Latency is reportedly about 70–500 milliseconds; input costs about $0.042 per million tokens and output is free. Versus text reasoning that often takes seconds, that latency sits closer to on-chain block cadence — for example Monad is reportedly about 300 milliseconds per block — than to microsecond co-lo races in U.S. equities. One existence proof is Jarrod Watts’s open-source jev-trader: according to public materials, it reads the MON-USDC book on Kuru on Monad, calls Jev each block, and posts real post-only limits; the repo quickly drew many stars, and decision latency has a public figure around 81 milliseconds. Those numbers come from public materials and should not be extrapolated straight into your environment.

The core claim compresses to one line: Jev is a decision engine, not the trading system itself. The state engine, strategy gates, risk vetoes and 24/7 loop still have to be built to production quality by you. On the scaffolding side, tools such as AgenKit can help land multi-layer scaffolding faster; what decides whether you can go live is still which layer you hand to the model — and which layer you never do.

Deterministic Code Owns State; The Model Only Judges

A common failure mode is asking: “Look at BTC and tell me what to do.” That outsources the whole trading system to one fuzzy prompt. A stabler split is:

Code computes state → Jev interprets state → code applies strategy → the execution layer places orders.

Anything you can compute stays in code: mid, spread, imbalance, realized volatility, inventory, drawdown, VWAP, queue position — do not outsource arithmetic to the model. Ask Jev only when judgment is needed: is the regime trend or mean reversion? Does this flow look informed or noisy? Is the quoting environment worth making markets in? Has execution already degraded? The official methodology stresses the same split: deterministic facts in code; fuzzy judgment to the model. Keep questions atomic and combine weights in code; when priorities change, retune coefficients instead of rewriting a whole prompt.

Split screen: code modules compute book state on the left, a decision engine returns probabilities and confidence on typed questions on the right, with a strategy-gate arrow in the middle

Three Primitives and Parallel Batteries

According to docs.typesafe.ai, Jev exposes three primitives you can mix in one call:

  • Noul: returns a 0–1 scalar (for example “flow toxicity” → 0.83).
  • Choice: picks among up to 255 options and returns a distribution (for example regime → trending 0.63).
  • Score: scores on a scale you define (for example quoting environment 0–3 → 2.3).

Each question is evaluated in parallel and in isolation; asking more usually does not make latency worse in a linear way. Reportedly it does not generate token-by-token autoregressively; it computes probability distributions in a parallel process. Training materials mention RLCD calibrated against real outcomes, so higher confidence is more checkable in aggregate. Context is about 32K tokens — enough for a compact snapshot and recent trades, and not a place to dump a full strategy manual. Narrow questions, tight state; everything else belongs in code.

In some of TypeSafe’s own workflow evals, public comparisons show agreement rates close to frontier models at much lower unit cost; those are vendor figures and only a directional reference. The cost story likewise comes from public materials: a full-load decision every block, year-round, can sit in the tens of dollars per month; hard-running the same loop on a frontier chat model can land an order of magnitude higher on the bill.

From Waitlist to the First Typed Decision

Follow the official quickstart from zero: apply for the waitlist at typesafe.ai → create a key → install the SDK (Python 3.10+) → fire the first decision. For higher frequency, public materials suggest preferring a direct path to https://api.typesafe.ai/v1/systemone to skip a gateway hop; pin the model version and log it on every response — confidence gates are calibrated to a specific version, and a silent upgrade can quietly break the system. Exact fields and SDK usage follow the docs.

The state engine is the layer most people under-invest in. Each block should emit a dense numeric snapshot kept under about 400 tokens when possible: you only pay for input tokens; timestamp discipline is absolute (fields may use only information available before the decision); and keep a triple log of snapshot–decision–outcome for later calibration. Professional mode is not one question at a time — it packs a parallel battery in one call (regime, toxicity, whether quoting is worth it, whether execution has degraded, and so on), then your strategy engine digests the answers with code thresholds. Thresholds live on your side: set confidence floors per action, higher when a wrong call costs more; fractional Kelly–style sizing only holds once probabilities are truly calibrated.

The 24/7 Loop and Hard Risk You Never Outsource

Half of market making is decades of deterministic math — for example Avellaneda–Stoikov reservation price and spread, computed in code every block. Jev only answers the part the formula cannot: is it worth quoting into this moment? A full loop usually splits into testable modules: read the book, compute state, parallel judgment, strategy gate, pricing, order, confirm, ledger, degrade. Two details decide whether it survives a live book:

  • Block-deadline rule: if the Jev decision does not return before the next block, hold or stand down — never quote from stale state. jev-trader has publicly handled the same case.
  • Gas honesty: a naive cancel-and-replace every block can lose structurally on fees; you may need wider spreads, longer resting time, or a real directional edge from the judgment battery. Run the P&L math before going live.

The risk engine never goes to Jev. Hard-coded vetoes check before every order: inventory caps, daily loss circuit breakers, link health, clock drift, and so on — using verifiable signals outside the model (is the file there, is a metric below a threshold), not a line that says “the system said it passed.” Write the degrade ladder in stone too: on model timeout, feed cut or exchange reject, how you shrink size, flatten only, then safe halt. Without that layer, “24/7” often means “until 2 a.m.”

A circular 24/7 loop lights up in turn — read book, state, judgment, strategy, orders and hard risk vetoes — with a block-deadline clock beside it

Where It Fits, and Honest Limits

The more realistic landing zones today are block-cadence on-chain order books (Kuru on Monad, some Solana DEXes, venues like Hyperliquid) and prediction markets where spreads often sit in hundreds of basis points. Microsecond co-lo races on top U.S. equities are a different infrastructure stack; bringing a 300-millisecond loop into a microsecond fight is the wrong match.

Calibrate before you worship a backtest. Do not only ask whether Jev guessed price; test the whole strategy under the same data, costs and limits against hand-written rules, a frontier LLM layer, a Jev layer, and “Jev + confidence gates.” Watch Sharpe, drawdown, hit rate, slippage, adverse selection, cost per million decisions and coverage. The real research question is often: does abstaining when uncertain make the ledger better? Calibration itself needs a reliability curve — when the model says 0.80, is the empirical frequency near 80% — checked with Brier, log loss, ECE and the like; if your distribution differs from the vendor training set, do a Platt-style recalibration in the strategy layer.

According to public materials, Jev can roughly: return typed, calibrated decisions in about 70–500 milliseconds; evaluate many questions in parallel in one call; take about 32K of state and up to 255 options; sit inside an about-300-millisecond block loop; replace some fuzzy rules, classifiers and regime heuristics; support sizing rules once you have verified calibration; and make “a battery of judgments every block” affordable.

What it cannot do: write a strategy manual or long prose; chain dependent reasoning inside one call (questions are isolated; dependencies need a second request); win a microsecond race; replace market-data infrastructure, numeric compute, exchange connectivity or hard risk; guarantee vendor numbers under your load; turn a losing strategy into a winning one. Decisions can get cheaper and faster — the edge is still your work.

Split the system into “deterministic compute + probabilistic judgment,” then check whether a calibrated decision model can carry the second half — that is closer to engineering than chasing the newest chat model. Related entry points: TypeSafe, docs, AgenKit, Monad, Kuru, jev-trader.

This article is for informational purposes only and does not constitute investment advice. Trading and smart-contract risk are your own.