
One Model, Many Conversations: LLM Serving, Memory, and Reasoning
How serving engines share one model across users, isolate KV caches, batch token generation, preserve context, and spend more inference compute on reasoning.
Read article →The field journal
Practical architecture decisions, tradeoffs, and operating lessons from software that has to survive business change.
Browse all topics →Showing 5 of 52 articles
Latest writing

How serving engines share one model across users, isolate KV caches, batch token generation, preserve context, and spend more inference compute on reasoning.
Read article →Trace llama_decode through GGML graphs, CUDA kernels, GPU memory, and next-token generation—and see why expensive LLM inference can still be fast.
Follow an LLM request past the API and into GGUF tensors, learned parameters, attention, and the transformer layer that turns model data into logits.
How a silent 10,000-record cap, ambiguous identities, and competing sources of truth changed a production membership reconciliation.
Why durable business capabilities should define platform boundaries, while commerce, payment, and ERP vendors remain replaceable implementation details.