Memory Is Not Context: Token Budgets, Narrative Relations, and Agent Economics

An agent answered a question about Infest with 2,326 tokens of memory in its final context. Getting to that answer required thirteen model calls and 19,831 input and output tokens. The verifier accepted it. One of its claims attributed evidence from the fourth song to the first. That execution contains the central problem of this study. A smaller context can conceal a more expensive decision process, and an accepted answer can conceal an unsupported interpretation. Neither becomes visible if the experiment ends at the final prompt or the aggregate score. ...

September 11, 2026 · 19 min · Updated September 11, 2026

What Should an Agent Remember? Album Narratives with LangGraph and MLflow

An album ends. An agent has read every lyric. We ask it whether the ending transforms something established near the beginning. The agent can produce a convincing paragraph with almost no memory. That is precisely the problem. Fluency gives us very little evidence that it remembered the right songs, preserved their differences, or resisted turning a collection of voices into one convenient story. In my previous post on MLflow, I argued that a prompt belongs in the release process because changing an instruction changes application behavior. Here I want to bring that argument into my research on musical narratives. Memory selection changes behavior too. It needs an evaluation contract. ...

September 11, 2026 · 19 min

MLflow for AI Engineering: Prompts Are Release Artifacts

A prompt can be three paragraphs long and still change the behavior of an entire application. It can decide whether an assistant answers, abstains, calls a tool, or invents the argument that the tool receives. Yet it is often reviewed with less discipline than a change to a configuration file. Let’s start there. If editing a sentence can change what your software does, that sentence belongs in your release process. ...

September 8, 2026 · 25 min · Updated September 9, 2026

AI Architecture - Notions on Training and Inference

CPU · GPU · TPU · Edge Computing The problem is not that AI is expensive. It’s that for years, people paid to train models as if that were the main cost — when the real cost, the one that never stops, is serving every response. TL;DR Inference costs exceed training by 15x–20x over a model’s operational lifetime. Optimizing for training while ignoring inference is optimizing the wrong problem. CPU (Intel Xeon AMX): the correct choice when the model lives alongside the data. Network latency kills any compute gain from moving to a GPU cluster. NVIDIA GPU (Blackwell/Hopper + TensorRT-LLM): still the default for research and heterogeneous production. CUDA is a 20-year moat. Don’t lock in at peak prices. Google TPU v6/v7: the right answer for high-volume, predictable inference. Midjourney cut monthly costs from $2.1M to $700K. The CUDA migration barrier no longer exists. Edge AI: thermodynamics, not algorithms, sets the limits. Pi 5 + Hailo-10H delivers 320 ms TTFT (6.4× faster than CPU-only) with a PCIe x1 bottleneck you need to design around. The right hardware is not the most powerful. It is the one that matches the problem topology to the silicon architecture without wasting energy or budget. Introduction In 2023, Nvidia published a post titled What Is AI Computing? focused on handling intensive computations — particularly useful for embedding design and optimization processes in Machine Learning — and advancing toward hardware acceleration to find patterns in immense amounts of data, thereby updating the assumptions of ML or AI models. All of this typically runs on GPUs. ...

April 5, 2026 · 15 min

Anatomy of an MLOps Pipeline - Part 1: Pipeline and Orchestration

Complete MLOps Series: Part 1 (current) | Part 2: Deployment → | Part 3: Production → Anatomy of an MLOps Pipeline - Part 1: Pipeline and Orchestration Why This Post Is Not Another Scikit-Learn Tutorial Most MLOps posts teach you how to train a Random Forest in a notebook and tell you “now put it in production.” This post assumes you already know how to train models. What you probably don’t know is how to build a system where: ...

January 13, 2026 · 29 min