Memory Is Not Context: Token Budgets, Narrative Relations, and Agent Economics

An agent answered a question about Infest with 2,326 tokens of memory in its final context. Getting to that answer required thirteen model calls and 19,831 input and output tokens. The verifier accepted it. One of its claims attributed evidence from the fourth song to the first. That execution contains the central problem of this study. A smaller context can conceal a more expensive decision process, and an accepted answer can conceal an unsupported interpretation. Neither becomes visible if the experiment ends at the final prompt or the aggregate score. ...

September 11, 2026 · 19 min

What Should an Agent Remember? Album Narratives with LangGraph and MLflow

An album ends. An agent has read every lyric. We ask it whether the ending transforms something established near the beginning. The agent can produce a convincing paragraph with almost no memory. That is precisely the problem. Fluency gives us very little evidence that it remembered the right songs, preserved their differences, or resisted turning a collection of voices into one convenient story. In my previous post on MLflow, I argued that a prompt belongs in the release process because changing an instruction changes application behavior. Here I want to bring that argument into my research on musical narratives. Memory selection changes behavior too. It needs an evaluation contract. ...

September 11, 2026 · 19 min