Does your AI answer have the evidence it needs?
I help teams examine how their AI applications retrieve information, remember it, and turn it into an answer. The starting point is a concrete failure: a plausible response with weak support, an agent that loses a key distinction, or a memory policy whose cost is difficult to justify.
I bring experience in AI engineering, NLP, statistics, and operating systems in production. My research on musical narratives explores related measurement problems; the work with your team starts from your application and its own evidence.
A memory and evidence diagnostic
We agree on one workflow and a small set of representative cases before starting. A typical starting scope is 20–30 cases, reviewed over two calendar weeks. Timing and price are agreed after checking access, complexity, and the work required.
| You bring | We examine | You receive |
|---|---|---|
| A workflow and examples of answers that concern you | Retrieval, source attribution, support, and missing evidence | An error analysis tied to specific cases |
| Traces or outputs and source material you are authorized to share | The context and calls behind the answer | A resource account where the available traces support it |
| A decision your team needs to make | One practical baseline or alternative within the agreed scope | A comparison, a reusable rubric, and prioritized recommendations |
The engagement ends with a written report and a 60-minute working session. We distinguish what the evidence establishes, what remains uncertain, and which experiment should come next. Improvements in quality or cost are outcomes to test, not guarantees.
When this is a useful fit
- You have a prototype or an existing application and need a clearer evaluation process.
- Your system retrieves relevant-looking material but sometimes draws unsupported conclusions.
- You want to understand whether extra memory, retries, or agent steps earn their cost.
Production integrations, large-scale data preparation, and ongoing implementation are scoped separately. If you need help defining the evaluation question first, describe that in your message.
See how I approach the work
- Memory Is Not Context — a research case about resource accounting and a fallible verifier.
- Prompts Are Release Artifacts — a proposed release workflow using a synthetic documentation example.
- Experiments and materials — code, recorded outputs, review material, and the limits of each study.
These are public research and educational examples. They are not client case studies or evidence of production improvements.
Tell me what your system is getting wrong
Tell me what you are building, one failure you want to understand, and the decision you need to make. Start with a non-sensitive description; we can agree on appropriate access before reviewing data.
You can also find me on LinkedIn. For my background, read About.