Does your AI answer have the evidence it needs?

I help teams examine how their AI applications retrieve information, remember it, and turn it into an answer. The starting point is a concrete failure: a plausible response with weak support, an agent that loses a key distinction, or a memory policy whose cost is difficult to justify.

I bring experience in AI engineering, NLP, statistics, and operating systems in production. My research on musical narratives explores related measurement problems; the work with your team starts from your application and its own evidence.

A memory and evidence diagnostic

We agree on one workflow and a small set of representative cases before starting. A typical starting scope is 20–30 cases, reviewed over two calendar weeks. Timing and price are agreed after checking access, complexity, and the work required.

You bringWe examineYou receive
A workflow and examples of answers that concern youRetrieval, source attribution, support, and missing evidenceAn error analysis tied to specific cases
Traces or outputs and source material you are authorized to shareThe context and calls behind the answerA resource account where the available traces support it
A decision your team needs to makeOne practical baseline or alternative within the agreed scopeA comparison, a reusable rubric, and prioritized recommendations

The engagement ends with a written report and a 60-minute working session. We distinguish what the evidence establishes, what remains uncertain, and which experiment should come next. Improvements in quality or cost are outcomes to test, not guarantees.

When this is a useful fit

  • You have a prototype or an existing application and need a clearer evaluation process.
  • Your system retrieves relevant-looking material but sometimes draws unsupported conclusions.
  • You want to understand whether extra memory, retries, or agent steps earn their cost.

Production integrations, large-scale data preparation, and ongoing implementation are scoped separately. If you need help defining the evaluation question first, describe that in your message.

See how I approach the work

These are public research and educational examples. They are not client case studies or evidence of production improvements.

Tell me what your system is getting wrong

Email me about a diagnostic

Tell me what you are building, one failure you want to understand, and the decision you need to make. Start with a non-sensitive description; we can agree on appropriate access before reviewing data.

You can also find me on LinkedIn. For my background, read About.