BlogEngineering
When our music agent said it changed a clef it never touched
An edit that changed nothing came back as Applied, the score reads hid a mid-piece clef, and the model believed the picture. Here is each cause and the fix, and what we have not measured yet.
The short answer
On October 10, 2026, Starling, our in-app assistant, told someone it had changed the left hand of bar 4 to bass clef. The bar was still in treble clef. Three things lined up: the edit tool answered “Applied” to an edit that changed nothing, our text reading of the score did not show clef changes inside the piece, and the model trusted a picture of the page. We fixed all three in the tools, then added one shared rule; the before-and-after measurement has not been run yet.
What happened?
A piano score had its lower staff switch to treble clef for one bar, bar 4, where the broken chord sits high, and back to bass clef after it. The person asked Starling, beside the score, to put bar 4’s left hand in bass clef. Starling called the edit tool, got a success back, and said it was done. The page still showed treble clef in bar 4. When the person said nothing had changed, the model agreed with whichever account seemed more likely, rather than reading the bar again.
We found it by reading the saved chat and the score’s version history on our staging server, the same day.
Why did the assistant believe the edit worked?
Because every source it had said so, or said nothing. Each cause on its own would have been survivable.
- Ask“bar 4, left hand, bass clef”
- Editsets the clef the staff opens with, which was already bass
- Result“Applied”, with a new version number
- Readshows the opening clef only, so nothing looks wrong
- Reply“Done”
- The edit changed nothing and still succeeded. Our tools treated any valid request as a new version. Setting a staff to the clef it already had was valid, so the server saved an identical score and replied “Applied” with a new version number. The model had no reason to doubt it.
- The reading hid what mattered. The text reading of a score listed each part’s opening clef and key, not the changes inside the piece. A clef that switched in bar 4 was invisible in text. Worse, a courtesy clef at the end of a bar, printed to warn of the next bar’s clef, was counted as that bar’s own clef.
- The model trusted the picture. When the text and a rendered page disagreed, or the text said nothing, the model judged from the image, where small clefs inside a system are easy to misread.
What did we change?
Four changes, all made on October 10, 2026, in the server that Starling and the MCP connector share.
- An edit that leaves the score as it was is refused. The tool returns an error, “Nothing changed: this edit leaves the score as it was, so no new version was saved”, and the version number stays put. Undo and redo are exempt.
- Every edit lists what it wrote. The result now carries lines such as
bar 4: lower staff clef treble -> bassorbars 1-4: key C major -> F major, so the model reports what the score says, not what it asked for. - Reads show changes inside the piece. The score reading gives each bar’s clefs, key and meter where they change, the review text marks “treble clef from here” and key changes in place, and a courtesy clef now counts toward the bar it announces.
- One shared rule for every route. The instructions for Starling, the MCP server and our plugin’s skills now say: say only what tool results show; an error or “Nothing changed” means the edit was not made, so say so; judge clefs, keys and notes from the text reading, not a picture; and if the person disputes it, read those bars again.
The order matters. The first two put the facts into the result the model reads, which no wording in a prompt can do. MCP’s specification describes the same split: errors from running a tool belong in the tool result, marked isError, so the model can see them and correct itself. Our “Applied” was a success where an error should have been.
How do we test it?
Two ways, one finished and one waiting.
- A check in our test suite opens a score whose lower staff switches clef inside a bar and whose key changes later, then confirms that every read shows both changes, that an edit lists the clef and key it wrote, and that an edit changing nothing is refused without a new version. It runs with our other quick checks on every pull request.
- An evaluation sends Starling six requests about a twelve-bar piano piece built like the one that went wrong: a real clef fix, the same fix on a bar already in bass clef, a question about the key after the change, a pitch edit that changes nothing, a pedal edit and a bar count. A separate model lists each claim in the reply and marks it against the score as saved after the turn, never against the assistant’s own words. We have not run it before and after the change yet, so we have no hallucination rate to report.
Rules lower the rate of false claims; they do not make it zero. A model can still misread a bar from the picture alone, which is why the rule sends it to the text and why the edits it makes are marked and undoable.
What would we tell someone building an agent with tools?
- A tool should never report success for nothing. If a call cannot change state, return an error the model can read.
- Return what was written, in the domain’s own terms, rather than echoing the request.
- Put every fact the model needs to check its work in text. If the facts are only in a picture, the model will judge from the picture.
- Read the conversations. We found this from one saved chat, not from a metric.
Sources
- Model Context Protocol, Specification 2025-11-25: Tools, section Error Handling.
- ScoreStarling’s saved chat and score history from October 10, 2026, and the change and tests that followed.
Questions and answers
Why would a tool report success for an edit that changed nothing?
Ours treated every valid request as a new version: setting a clef to the clef it already had was valid, so it saved an identical score and replied Applied. Since October 10, 2026 such an edit is refused with “Nothing changed”, and no version is saved.
Do prompt rules alone stop an agent claiming edits it did not make?
Not reliably. A rule helps, but the model still has to know what is true. The changes that matter most are in the tools: refusing no-op edits, and listing what each edit wrote, so the facts are in the result the model reads.
Does this apply to ChatGPT and Claude too?
Yes. Starling and the MCP connector call the same tools, so ChatGPT, Claude and other MCP clients get the same refusals and the same lists of written changes, and the same rule is in the server’s instructions.