Quoters
One contract, three quoters
The same service written three times, in Go, TypeScript and Python. Each takes a cart in free text and answers a priced quote or a refusal, through the same stages, the same prompts and the same rules.
Overview
A quoter is one stateless HTTP service: a cart goes in (POST /v1/quotes), a quote comes out, or a 422 problem saying which stage refused the cart. Inside, it runs the pipeline of the Architecture page: prepare, guard, parse beside recount, identify, judge, price. Models read; code counts and prices. The web app's BFF talks to one quoter at a time, and any of the three will do.
Why three. The implementations are compared on measurements (accuracy, latency, cost, from the system bench of e2e/), not on opinions. Holding three of them to one suite also shows where the contract or the shared rules are ambiguous: a rule two languages read differently gets written down.
What they share.
- The contract, api/openapi.yaml: routes, bodies, problem codes, headers. Each quoter derives its types from it in its own way (below).
- The prompts,
prompts/: every word put to a model, one JSON file per stage./healthzreports each file's version (the first 8 hex digits of its SHA-256), and the suite fails a quoter that runs other prompts. - The fake engines (
ENGINES=fake): deterministic stand-ins for Jev and the LLMs, with the same rules in the three, driven by#fake:lines in the cart. No key, no cost. - The end-to-end suite in
e2e/, black box, run against each quoter on the fake engines (task e2e:all), and the shared cases. - The trace shape in Langfuse (Usage, cost and traces): one trace named
quoteper request, anagentroot taggedquoter:…andengines:…, with the user, the session and the cart as prepared for input and the response body as output; one observation per stage, typed for Langfuse's graph (guarda guardrail;parse,recount,identifychains;judgean evaluator); one generation per model call with the cost OpenRouter billed; and four scores on the trace,cost_usd,latency_ms,attemptsandoutcome. A dashboard reads the three alike.
Each one listens on its own port. The buttons open its /healthz, which names the implementation, the engines, the tracing and the prompt versions.
Differences, side by side
Every cell is read from the quoter's code, manifest or Dockerfile. Rows with an amber edge are where the three made different choices of design, not only of library.
| Go | TypeScript | Python | |
|---|---|---|---|
| Language, runtime | Go 1.26.4 (go.mod), one static binary |
TypeScript 5.9 on Node 26 (engines: >=26); Node runs the .ts sources by type stripping, no build step |
Python 3.14 (.python-version), dependencies locked by uv |
| HTTP layer | net/http and its ServeMux, wrapped by the generated HandlerWithOptions; log/slog |
Hono 4.13 on @hono/node-server 2.1 |
FastAPI 0.142 on Starlette 1.7, served by uvicorn 0.54; FastAPI's validation and generated OpenAPI are off |
| From the contract | Generated server: oapi-codegen writes the ServerInterface and the models into internal/httpapi/openapi.gen.go |
Generated types: openapi-typescript writes src/generated/openapi.ts; the Hono routes are written by hand |
Hand-written pydantic models (api/models.py); tests/test_contract.py holds their fields, required fields and enums to api/openapi.yaml |
| LLM client (parse, recount) | trpc-agent-go 1.11.2's OpenAI model (model/openai, on openai-go 1.12): one GenerateContent call, no agent, no tool; the answer decoded strictly by hand (decodeReading) |
openai 7.27 SDK, maxRetries: 0; answer checked with Ajv 8 against the schema of parse.json |
openai 3.24 AsyncOpenAI, max_retries=0; answer validated by pydantic |
| Jev client | Hand-written over net/http (internal/decide): HTTP.Decide, DecideAll, WithRetry for the benches |
Hand-written class Jev over fetch: decide, decideAll, retries only when attempts > 1 |
Hand-written class Jev over httpx.AsyncClient, the wire checked by pydantic: decide, decide_all |
| Tokenizer | o200k_base, ranks embedded by tiktoken-go-loader, split with regexp2; its own BPE merge on a heap, O(n log n), held to tiktoken-go's counts by a test | o200k_base, ranks shipped by js-tiktoken; its own BPE merge on a heap, held to js-tiktoken's counts by a test | tiktoken 0.14, encode_ordinary (Rust BPE); the vocabulary fetched once, at setup or in the Docker build, checked against a pinned SHA-256, never at run time |
| Prompts | Compiled in: prompts/ is a small Go module (//go:embed *.json), required through a replace to its path; read and checked at init |
Read at startup: loadPrompts reads the four files from PROMPTS_DIR (the repository's prompts/ by default) and checks them |
Read at startup: load_prompts reads the four files from PROMPTS_DIR into pydantic models and checks them |
| Concurrency | Goroutines: sync.WaitGroup.Go for parse beside recount, errgroup with SetLimit(16) for Jev; context.Context cancels |
One event loop: Promise.allSettled for parse beside recount, a pool of 16 async workers for Jev; AbortSignal cancels |
asyncio: create_task for the recount, TaskGroup under a Semaphore(16) for Jev; task cancellation |
| Tracing SDK | trpc-agent-go's telemetry/langfuse (OpenTelemetry, OTLP) for the spans, opened with its tracer, a no-op until started; the scores posted to Langfuse's ingestion API by a queue of its own |
@langfuse/otel span processor on OpenTelemetry's NodeTracerProvider; observations with @langfuse/tracing; scores with @langfuse/client |
langfuse 4.16 SDK on a TracerProvider of its own (OpenTelemetry SDK 1.45), behind a Tracer protocol; scores with its create_score |
| Tracing | The pipeline makes the trace: Pipeline.Quote opens the quote root and ends it with the body the handler builds through Request.Answer; telemetry.Scores queues the scores off the request's path |
The HTTP layer makes the trace: traced in app.ts opens the root around pipeline.quote, which opens only the stage spans, then sets the body and sends the scores |
The pipeline makes the trace: Pipeline.quote opens the root through its Tracer, takes the body from Request.respond, and _measure scores it |
| Lint, types | go vet; golangci-lint 2, standard linters and 9 more; gofmt, goimports |
ESLint 10 with typescript-eslint strictTypeChecked; tsc strict, with noUncheckedIndexedAccess and exactOptionalPropertyTypes; Prettier |
Ruff 0.16, 22 rule sets, and its formatter; mypy 2.4 --strict with the pydantic plugin |
| Tests | go test -race: 99 test functions, 300 tests with their subtests bench harness apart: 20 more functions |
Vitest 5.0: 281 tests in 9 files | pytest 9.1 with pytest-asyncio: 377 tests, parametrized cases counted, in 12 files |
| Docker image | golang:1.26-alpine build, gcr.io/distroless/static-debian12:nonroot run: 49.2 MB |
node:26-alpine, production dependencies only: 374 MB |
python:3.14-slim with a uv 0.12 virtual environment and the vocabulary: 324 MB |
| Throughput under load | 200 in flight on fake engines that wait as models do: 57.6 req/s on 1 CPU at 7 %; with 10 ms of CPU per call, 26 req/s on 1 CPU, 56 req/s on 4 (253 % CPU) | The same: 57.0 req/s on 1 CPU at 17 %; with 10 ms of CPU per call, 20 req/s on 1 CPU and 20 on 4: one event loop, one core | The same: 57.4 req/s on 1 CPU at 14 %; with 10 ms of CPU per call, 18 req/s on 1 CPU and 19 on 4: one interpreter, one core |
| Latency under load | p50 3.5 s, p99 7.7 s at 200 in flight; CPU-bound on 4 CPUs, unchanged (3.6 s, 7.8 s) | p50 3.6 s, p99 7.7 s; CPU-bound on 4 CPUs, p50 6.7 s, p99 18.8 s | p50 3.6 s, p99 7.7 s; CPU-bound on 4 CPUs, p50 10.1 s, p99 22.1 s |
| Memory | 26 MiB at rest, 45 to 66 MiB at 200 in flight | 138 MiB at rest, 151 to 162 MiB | 141 MiB at rest, 148 to 166 MiB |
| Lines of code | 3 512 source, 3 498 test apart: 776 generated; bench harness 2 062 source, 845 test | 2 804 source, 2 211 test apart: 454 generated | 3 122 source, 2 734 test |
How these were measured, on 2 October 2026, with task ci green and every test passing. Lines: non-blank lines, comments included (grep -cv '^\s*$') over *.go, *.ts or *.py; source is cmd/delorean and internal/ for Go, src/ otherwise; tests are *_test.go, test/, tests/. Tests: tests run, as each runner counts them (go test -json, Vitest's JSON report, pytest). Images: docker images for delorean-quoter-go, -typescript and -python as Docker Compose builds them. All three move with the code. Load: task bench:load on 3 October 2026, each image alone with --cpus 1 or 4 and 512 MB, fake engines with FAKE_LATENCY=real, 200 clients in a closed loop; the method and every step are in Testing, Load bench.
Code structure
One section per quoter: a trimmed tree, then thirteen key functions in the same order in the three, so that a step reads across. Links open the file on GitHub at the function's line, on main.
Go
quoters/go/
├── cmd/delorean/main.go serve | version: config, telemetry, engines, http.Server
├── cmd/bench/ the component benches' command, not the service
└── internal/
├── httpapi/ the contract's surface
│ ├── openapi.gen.go generated by oapi-codegen: ServerInterface, models
│ ├── quotes.go POST /v1/quotes: headers, strict body, pipeline, answer
│ ├── httpapi.go New, /healthz, /v1/catalog
│ ├── middleware.go X-Request-Id, a log line per request, 404 and 405
│ └── problem.go RFC 9457 problems
├── prepare/prepare.go Normalize, the o200k_base Counter
├── pipeline/ the order of the stages, every rule that decides
│ ├── pipeline.go Quote, read, Merge, Identify, Recounted
│ ├── reading.go readAgain, readTwice, judge and its memo
│ ├── ports.go the engines' interfaces, Weigh
│ ├── rejection.go Rejection and its codes
│ └── trace.go the quote's trace, the stages' typed spans
├── live/ engines on OpenRouter
│ ├── parse.go Parser: one call to trpc-agent-go's OpenAI model
│ ├── guard.go, identify.go, judge.go the Jev questions
│ └── prompts.go the embedded prompts/, checked and versioned
├── decide/ the Jev client: HTTP, DecideAll, WithRetry
├── fake/fake.go ENGINES=fake
├── pricing/pricing.go Catalog.Price, in cents
├── telemetry/ the Langfuse exporter (trpc-agent-go), Scores
├── config/, cart/ the environment; Film, Mention, Line
└── bench/ the bench harness: subjects, runs, reports
-
Request handlerinternal/httpapi/quotes.go
func (s *server) CreateQuote(w http.ResponseWriter, r *http.Request, params CreateQuoteParams)The generated
ServerInterface's method. Checks the three headers, decodes the body strictly withreadCart(valid UTF-8, one object, no unknown field), and gives the pipeline acontext.WithTimeoutand anAnswercallback:s.answerbuilds the response (200, a*pipeline.Rejection422,ErrEngine502) inside the trace, which keeps it as its output. -
Prepareinternal/prepare/prepare.go
func Normalize(text string) stringLF line ends; control and format characters dropped but
\n,\tand the two joiners; NFC withgolang.org/x/text; trimmed. -
Token countinternal/prepare/prepare.go
func (c *Counter) Count(text string) intSplits with the o200k_base pattern (regexp2, which has the lookahead it needs) and counts each piece with
tokens: tiktoken's merge order, kept in a typed binary heap over a linked list of parts. -
func (p *Pipeline) Quote(ctx context.Context, req Request) (Quote, error)Opens the trace (
startTrace), runsread(prepare, guard, the read-again loop, price), fills the report, ends the trace with the answer and hands the measures toMeasured. A refusal comes back as the error, a*Rejectioncarrying its report; an engine failure wrapsErrEngine. -
Read-again loopinternal/pipeline/reading.go
func (p *Pipeline) readAgain(ctx context.Context, r *run, text string) (Reading, error)Up to
ReadAttemptsreadings:readTwice, identify, judge. Stops at the first judgement at the threshold or above; otherwise the next parse gets aRetrywith its raw reading and the failing findings. -
Guard decisioninternal/pipeline/ports.go
func Weigh(q GuardQuestions) GuardVerdictThe verdict of Jev's two answers: injection = steer, valid = (1 − steer) × order, invalid = (1 − steer) × (1 − order); a tie goes to the refusal. The engine calls it (
live.Guard.Check);readthen refuses belowGuardMinConfidencethroughguardRejection. -
Parse and recountinternal/pipeline/reading.go
func (p *Pipeline) readTwice(ctx context.Context, r *run, text string, attempt int, again *Retry) (raw, reading, recount []cart.Mention, err error)Two
wg.Gogoroutines: the parser and the blind recounter, each alive.Parser.Parse(oneGenerateContentcall on trpc-agent-go's OpenAI model). A parse that fails cancels the recount;Mergerefuses a title over the copy limit. -
Identifyinternal/pipeline/pipeline.go
func Identify(ctx context.Context, id Identifier, known map[string]Identification, readings ...[]cart.Mention) ([][]cart.Line, Usage, error)Asks only the titles
knownlacks, from the reading and the recount in one call, then remembers them: no title is identified twice in a request. On Jev,live.Identifier.Identifysends one request per title. -
Judge and its memointernal/pipeline/reading.go
func (p *Pipeline) judge(ctx context.Context, r *run, n int, text string, judged map[string]Judgement, lines, recounted []cart.Line) (Judgement, error)Puts a reading to Jev (
live.Judge.Judge: asked and identity per line, missing for the whole) only whenjudgedlacks itsreadingKey; a reading seen before is reused throughinLineOrder.Recountedthen adds the count checks. -
Pricinginternal/pricing/pricing.go
func (c Catalog) Price(lines []cart.Line) QuoteUnit prices in integer cents, then the highest saga tier the distinct volumes reach, taken off the saga lines only, rounded half up.
-
Fake enginesinternal/fake/fake.go
func New() pipeline.EnginesA guard on injection marks, a reader of one mention per line, an identifier of the English titles, a judge driven by the
#fake:lines. -
Trace setupinternal/telemetry/telemetry.go
func Start(ctx context.Context) (shutdown func(context.Context) error, enabled bool, err error)Starts trpc-agent-go's Langfuse exporter when
LANGFUSE_*is set; until then its tracer is a no-op.mainthen setsPipeline.Measuredto atelemetry.NewScoresqueue. Each model call is a generation:Parser.callfor the LLMs,decide.HTTP.Decidefor Jev. -
Trace shapeinternal/pipeline/trace.go
func (p *Pipeline) startTrace(ctx context.Context, req Request) (context.Context, trace.Span)The
agentroot namedquote, taggedquoter:goandengines:…, with the request id, the prompt versions, the user and the session.startStagetypes each stage fromobservationTypes;endTracesets the body sent, the outcome, the attempts and the total;Scores.Quotequeues the four scores.
TypeScript
quoters/typescript/
├── src/
│ ├── main.ts serve | version: config, tracing, prompts, engines, server
│ ├── config.ts the environment, checked
│ ├── http/
│ │ ├── app.ts createApp: Hono routes, the quote's trace, problems
│ │ ├── body.ts readBody, decodeQuoteRequest
│ │ └── contract.ts the pipeline's values in the contract's types
│ ├── generated/openapi.ts generated by openapi-typescript
│ ├── prepare/
│ │ ├── normalize.ts normalize
│ │ └── tokens.ts TokenCounter, the heap merge
│ ├── pipeline/
│ │ ├── pipeline.ts Pipeline, Run, the read-again loop
│ │ ├── reading.ts verdictOf, merge, Identifications, countFindings
│ │ ├── ports.ts the engines' interfaces, EngineError
│ │ └── rejection.ts Rejection, Report
│ ├── engines/
│ │ ├── live/index.ts liveEngines
│ │ ├── live/jev.ts the Jev client, over fetch
│ │ ├── live/questions.ts jevGuard, jevIdentifier, jevJudge
│ │ ├── live/reader.ts llmReader: openai SDK, Ajv
│ │ └── fake.ts ENGINES=fake
│ ├── prompts.ts loadPrompts, promptVersions
│ ├── pricing.ts price, in cents
│ ├── telemetry/ langfuse.ts: startTracing, scores; trace.ts: observe
│ └── cart.ts, text.ts, log.ts
└── test/ Vitest, one file per area
-
Request handlersrc/http/app.ts
app.post('/v1/quotes', async (c) => …) // in createApp(config: AppConfig): Hono<Env>Checks the headers against
HEADER_FORMATS, reads the body withreadBodyanddecodeQuoteRequest, passesAbortSignal.anyof the client's signal and the budget's, and runsquoteinsidetraced: aRejectionis caught as 422, anEngineErroras 502;contract.tsmaps to the generated types. -
Preparesrc/prepare/normalize.ts
export function normalize(text: string): stringtoWellFormed()first, so a lone surrogate reads as U+FFFD as in Go, then the same steps: one Unicode-property regex,normalize('NFC'),trim(). -
Token countsrc/prepare/tokens.ts
TokenCounter.count(text: string): numberSplits with js-tiktoken's o200k_base pattern and counts each piece's UTF-8 bytes with
#countPiece: the same heap merge, a pair keyed as one number, rank × 2³² + start. -
Pipeline.quote(request: QuoteRequest, signal: AbortSignal): Promise<Quote>Runs
#read(prepare, guard, the loop) under the request's span, which the HTTP layer opened: the pipeline traces its stages, not the quote. A refusal is thrown, aRejectionwith its report attached, as is anEngineError. -
Read-again loopsrc/pipeline/pipeline.ts
#readUntilFaithful(run: Run, text: string): Promise<{ price: Price; judgement: Judgement }>The same loop, which also prices: the first reading that passes is priced inside it, and after the last attempt it throws
unfaithful_reading. -
Guard decisionsrc/pipeline/reading.ts
export function verdictOf({ order, steer }: GuardAnswers): GuardVerdictThe same weighing, on the pipeline's side: the engine (
jevGuard) returns the two raw answers.Pipeline.#pass(v)then lets a confident valid verdict through, or throws. -
Parse and recountsrc/pipeline/pipeline.ts
#readTwice(run: Run, text: string, retry: Retry | undefined, first: boolean): Promise<Reading>Promise.allSettledof the two stages; a parse that fails aborts the recount through anAbortController. Each reader is anllmReader: the openai SDK, the answer checked with Ajv. -
Identifysrc/pipeline/reading.ts
class Identifications { unknown(...readings) · learn(titles, identifications) · lines(mentions) }The request's memory of titles:
unknownlists what to ask,learnchecks and keeps Jev's answers,linesgives each mention its film. On Jev,jevIdentifier. -
Judge and its memosrc/engines/live/questions.ts
export function jevJudge(jev: Jev, prompts: Prompts['judge']): JudgeOne request per probe of
judgeProbes. The memo lives in#readUntilFaithful: aMap<string, Finding[]>byreadingKey, reordered byinLineOrder;countFindingsadds the count checks anew. -
Pricingsrc/pricing.ts
export function price(catalog: Catalog, lines: readonly Line[]): PriceThe same rules as Go's, written with
map,filterandreduce; cents in plain numbers, which stay exact at these sizes. -
Fake enginessrc/engines/fake.ts
export function fakeEngines(): EnginesThe same five stand-ins as plain objects; the
#fake:lines are named inDIRECTIVE. -
Trace setupsrc/telemetry/langfuse.ts
export function startTracing(env: Env): TracingRegisters a
NodeTracerProviderwith Langfuse's span processor whenLANGFUSE_*is set, and aLangfuseClientwhosescoresends the scores. Spans come fromobserve(name, type, fn, metadata?), found in the async context as Go finds them in acontext.Context. -
Trace shapesrc/http/app.ts
function traced(c: HonoContext<Env>, request: QuoteRequest, run: (traceId: string | undefined) => Promise<Answered>)Opens the root with
observe('quote', 'agent', …)andwithTraceAttributes, sets the tags, the metadata and the normalized cart, runs the quote, then sets the body as output and sends the four scores throughconfig.tracing.score. The stages' types areOBSERVATION_TYPES, in the pipeline.
Python
quoters/python/
├── src/delorean/
│ ├── __main__.py serve | version | tokenizer
│ ├── config.py Settings, from the environment
│ ├── api/
│ │ ├── app.py create_app, Service: the FastAPI routes
│ │ ├── body.py header, read_cart
│ │ ├── models.py the contract's bodies, pydantic, by hand
│ │ ├── answers.py the pipeline's values as bodies
│ │ └── problems.py, middleware.py RFC 9457 problems; request id, log line
│ ├── prepare.py normalize, TokenCounter (tiktoken)
│ ├── pipeline/
│ │ ├── pipeline.py Pipeline, _Run, the read-again loop
│ │ ├── rules.py guard_verdict, merge, count_findings, facts
│ │ └── ports.py, outcome.py the engines' protocols; Quote, Rejection, Report
│ ├── engines/
│ │ ├── __init__.py open_engines
│ │ ├── live/__init__.py open_live_engines: the httpx and OpenAI clients
│ │ ├── live/jev.py the Jev client, over httpx
│ │ ├── live/questions.py JevGuard, JevIdentifier, JevJudge
│ │ ├── live/reader.py LlmReader: openai SDK, pydantic
│ │ └── fake.py ENGINES=fake
│ ├── prompts.py load_prompts, pydantic models of the files
│ ├── pricing.py Catalog.price, in cents
│ ├── tasks.py all_of: a TaskGroup under a Semaphore
│ ├── telemetry.py Tracer, NoTracer, LangfuseTracer
│ └── cart.py, logs.py
└── tests/ pytest; test_contract.py holds models.py to the contract
-
Request handlersrc/delorean/api/app.py
async def Service.create_quote(self, request: Request) -> ResponseReads the headers with
headerand the body withread_cart(a pydanticQuoteRequest), runs the pipeline underasyncio.timeoutwith arespondcallback, so the trace's output is the body sent.Service.bodymatches the outcome:Quote200,Rejection422;EngineErrororTimeoutError502. -
Preparesrc/delorean/prepare.py
def normalize(text: str) -> strThe same steps, with
unicodedata.categoryto dropCcandCfandunicodedata.normalize("NFC", …). -
Token countsrc/delorean/prepare.py
def TokenCounter.count(self, text: str) -> intencode_ordinaryof tiktoken's o200k_base, whose Rust merge needs no rewrite.TokenCounter.loadrefuses a missing or altered vocabulary rather than download it. -
async def Pipeline.quote(self, request: Request) -> Quote | RejectionOpens the
quotetrace, with its tags and metadata, through theTracerit holds, and measures the outcome before it returns. A refusal is a value returned, not raised; only anEngineErrorraises. -
Read-again loopsrc/delorean/pipeline/pipeline.py
async def Pipeline._read(self, run: _Run, cart: str, trace: Trace) -> Quote | RejectionPrepare, guard, then the loop inline:
for attempt in range(1, self.read_attempts + 1), each failed attempt leaving aRetrywith the parser's answer and the failing findings. -
Guard decisionsrc/delorean/pipeline/rules.py
def guard_verdict(answers: GuardAnswers) -> GuardVerdictThe same weighing, among the pipeline's rules;
maxkeeps the first of equals, and the refusals come first.Pipeline._guard_rejectionbuilds the refusal. -
Parse and recountsrc/delorean/pipeline/pipeline.py
async def Pipeline._read_twice(self, run: _Run, text: str, retry: Retry | None) -> tuple[_Parsed, list[Mention]] | RejectionThe recount started with
asyncio.create_task, the parse awaited beside it, the recount cancelled infinally. Each reader is anLlmReader.read:AsyncOpenAI, the answer validated by pydantic. -
async def Pipeline._identify(self, run: _Run, memory: _Memory, parsed: list[Mention], recounted: list[Mention]) -> tuple[list[Line], list[Line]]Asks only the titles
memory.identifiedlacks, throughJevIdentifier.identify, then gives both readings their films withrules.lines. -
Judge and its memosrc/delorean/pipeline/pipeline.py
async def Pipeline._judge(self, run: _Run, memory: _Memory, text: str, reading: list[Line], recount: list[Line]) -> JudgementThe memo is
memory.judged, keyed byrules.facts(reading), a frozenset;JevJudge.judgeruns only on a miss, andrules.count_findingsanew. -
Pricingsrc/delorean/pricing.py
def Catalog.price(self, lines: Sequence[Line]) -> PriceThe same rules on frozen dataclasses; the discount in integer division,
(base * percent + 50) // 100. -
Fake enginessrc/delorean/engines/fake.py
def engines(pace: Pace = INSTANT) -> EnginesThe same stand-ins as classes, the recount a
FakeReader(recount=True). Picked byopen_engines, an async context manager that, for the live engines, owns the HTTP clients. -
Trace setupsrc/delorean/telemetry.py
class LangfuseTracer(*, public_key: str, secret_key: str, base_url: str, exporter: SpanExporter | None = None)The Langfuse SDK on a
TracerProviderof its own, leaving the global OpenTelemetry state alone;NoTracerstands in without configuration. The pipeline callstrace,spanandgeneration. -
Trace shapesrc/delorean/pipeline/pipeline.py
def _measure(self, trace: Trace, run: _Run, request: Request, outcome: Quote | Rejection | str) -> NoneWrites the outcome, the attempts and the total as the trace's metadata, and the four scores with
trace.score, which_RecordedTracesends ascreate_score, id<trace id>-<name>. The root is opened inPipeline.quote; the stages' kinds are_SPAN_KINDS.
One request, three ways
One cart, 2 x Back to the Future then Back to the Future Part III on the next line, accepted at the first reading. Each cell is the call chain that step takes; amber-edged rows are where the quoters part ways.
| Step | Go | TypeScript | Python |
|---|---|---|---|
| Route | generated HandlerWithOptions → server.CreateQuote |
Hono app.post('/v1/quotes') |
FastAPI route → Service.create_quote |
| Body | checkHeaders → readCart (encoding/json, DisallowUnknownFields) |
HEADER_FORMATS → readBody → decodeQuoteRequest |
header ×3 → read_cart (QuoteRequest.model_validate_json) |
| Budget | context.WithTimeout(r.Context(), RequestTimeout) |
AbortSignal.any([c.req.raw.signal, AbortSignal.timeout(…)]) |
async with asyncio.timeout(request_timeout) |
| Trace | in the pipeline: Pipeline.Quote → p.startTrace |
in the HTTP layer: traced → observe('quote', 'agent') → withTraceAttributes → pipeline.quote |
in the pipeline: Pipeline.quote → tracer.trace("quote", tags=…) |
| Prepare | prepare.Normalize → Counter.Count |
normalize → TokenCounter.count |
normalize → TokenCounter.count |
| Guard, 2 Jev requests | live.Guard.Check → decide.DecideAll → pipeline.Weigh, in the engine |
jevGuard.check → jev.decideAll; then verdictOf → #pass, in the pipeline |
JevGuard.check → jev.decide_all; then rules.guard_verdict, in the pipeline |
| Parse beside recount | readTwice: two wg.Go → Parser.Parse (trpc-agent-go model) → Merge / Tally |
#readTwice: Promise.allSettled → llmReader.read (openai SDK, Ajv) → merge |
_read_twice: create_task(_recount) + await _parse → LlmReader.read (openai SDK, pydantic) → rules.merge |
| Identify, 2 Jev requests | pipeline.Identify(known, …) → live.Identifier.Identify |
Identifications.unknown → jevIdentifier.identify → learn, lines |
_identify → JevIdentifier.identify → rules.lines |
| Judge, 5 Jev requests | p.judge → live.Judge.Judge → Recounted |
at.stage('judge') → jevJudge.judge → countFindings |
_judge → JevJudge.judge → rules.count_findings → rules.judgement |
| Price, 4 050 cents | Catalog.Price |
price(catalog, lines) |
Catalog.price |
| Answer | req.Answer → s.answer → s.quote(q), then endTrace; sent.send 200 |
pipeline.quote resolves → contract.quote → c.json 200, then traced sets the output |
request.respond → Service.body → answers.quote, then trace.output; _json 200 |
| Scores | p.Measured(m) → Scores.Quote, a queue → /api/public/ingestion |
config.tracing.score → LangfuseClient.score.create |
_measure → trace.score → create_score |
| Had it been refused | an error: in s.answer, errors.As(err, &rej) → s.rejected → problemResponse 422 |
an exception: in quote, catch, instanceof Rejection → contract.rejected 422 |
a value: in Service.body, case pipeline.Rejection() → answers.rejection → problem_response 422 |
Same stages, same calls to the models, same answer: the two saga volumes earn 10 % off the three DVDs (4 500 − 450 = 4 050 cents), after 9 Jev requests and 2 LLM calls in each quoter. They part ways on four things only:
- Where the trace is made. Go and Python open the
quotetrace in the pipeline and get the response body back through a callback; TypeScript opens it in the HTTP layer, around a pipeline that traces only its stages. The shape in Langfuse is the same. - Where the guard is weighed. Go's engine returns a verdict; in TypeScript and Python the engine returns the two raw answers, and the pipeline weighs them.
- How a refusal travels. An error value in Go, an exception in TypeScript, a returned value in Python. In all three, an engine failure is a different kind (
ErrEngine,EngineError) and answers 502. - How work runs side by side and stops. Goroutines under a
context.Context, promises under anAbortSignal, asyncio tasks under aTaskGroupand cancellation; the order in which a failure, a refusal and the recount decide is the same.