This publication brings interpretation and experiments together. A musical reading is an argument to support; a model score is a measurement whose meaning needs to be checked. This page records how to read the evidence and where published claims have changed.

Reading the evidence

  • Measured result: identify the corpus, configuration, comparison, and recorded output. Repeated model runs do not create new independent albums or questions.
  • Model judgment: a judge model supplies another fallible assessment. Agreement with it is not human validation.
  • Interpretation: an account of narrative, emotion, or intention must remain distinguishable from the numerical representation used to explore it.
  • Scope: a lyric-based analysis does not observe harmony, timbre, performance, or listener cognition. A small selected corpus cannot establish a universal limitation of an architecture.
  • Reproduction: inspecting a saved result, recomputing its aggregates, and rerunning the model are different actions. The experiment instructions specify which is possible and when API costs apply.

The experiment index links the available materials and current evidence status. When a numerical claim cannot be reconciled with its artifacts, it should remain unresolved rather than be promoted as a result.

September 27, 2026 — Attention Windows

The revised article replaces the title’s cognitive-load claim with a question about embedding similarity. It withdraws claims of structural impossibility, universal embedding failure, and effects on commercial recommendation systems that were not tested.

The earlier article acknowledged that a previous draft contained invented figures and stated that they had been replaced. That acknowledgment is retained here and in the article; this revision does not independently reconstruct that draft history.

A direct inspection of the existing attention_windows_results.csv found means of approximately 0.412 for 403 Beatles rows and 0.053 for 208 Pink Floyd rows. The previous headline reported 0.57 and 0.25. The export lacks enough run metadata to attribute that difference to a particular threshold or version. The earlier headline inference is therefore withdrawn pending reconciliation. This file inspection is not a new embedding experiment or scientific replication.

The revised article links a summary with the inspected file’s SHA-256 and the calculation. Model runs, human evaluation, and a reconciled inferential analysis remain outstanding.

September 27, 2026 — Aquamosh

The revised article narrows the claim to associations in the studied album and configurations. It retains descriptive break rates from the existing exports and identifies the comparison judge as GPT-4o-mini, not a human-equivalent reference.

The revision withdraws claims that the results falsify the distributional hypothesis, establish an architectural impossibility, or quantify failure in Spotify, moderation, or support systems. Those systems were not tested. Non-significant correlations in the audio extension are no longer interpreted as proof of independence or producer intention.

Aggregate tables were checked against existing files and packaged with hashes for inspection. Embeddings and judge calls were not rerun, and no new human ratings were collected. The previous long-form analyses remain in repository history; they are historical records, not the current interpretation.

Suggest a correction

Email a correction with the article, the claim or figure, and the evidence you would like me to examine. Please distinguish a reproduction failure from a different interpretation; both can be useful, but they require different checks.