BlogEngineering
Our MCP App in ChatGPT on mobile: what we learned
We ran our score panel in ChatGPT’s mobile website on an iPhone Simulator: what its safe-area insets contain, what happens when it rebuilds a panel, and how we now test both without a phone.
The short answer
On a phone, ChatGPT lays its title bar and message box over a fullscreen MCP App and reports them as safe-area insets, about 47 px and 127 px in our run, and moves the bottom one while the app is open. It also rebuilt our panel from the conversation’s original tool result whenever it entered or left fullscreen. Keep the inset clear without adding your own allowance, place script-positioned elements again on every context change, and check on load whether a result is out of date. We tested ChatGPT’s mobile website in an iPhone Simulator, not the native app.
What broke when we opened our MCP App in ChatGPT on a phone?
Four things, all involving fullscreen. ScoreStarling’s panel shows a score in the chat (how to connect it); in fullscreen you tap notes and ask the chat about them from a selection card that docks above the player on a phone. On October 7, 2026 we ran it in ChatGPT’s mobile website, in Safari on the iOS Simulator’s iPhone 17 (402 × 874 points). The panel’s frame is cross-origin, so we read it from screenshots.
| What we saw | Cause | Fix |
|---|---|---|
| The player floated about a third of the way up the screen | A fixed 96 px on top of a bottom inset (about 127 px) that already covers the message box | The inset or 96 px, whichever is larger |
| A longer draft or the Thinking bar raised the player, not the docked card | ChatGPT moves the inset; the player follows in CSS, the card is placed by script | Place the card on each context change |
| After an edit, entering or leaving fullscreen showed the old score, read-only | ChatGPT built a new panel from the original tool result | Read the current version if a result is behind |
| After a chat edit, the card floated about 40 px above the player | Our “Updated from the chat” notice moved the area the card is measured from | Place the card when that area resizes |
Why does a fullscreen MCP App hide behind ChatGPT’s message box on a phone?
Because ChatGPT draws its message box over your frame and tells you how much it covers through the bottom safe-area inset. Ignore the inset and your controls sit under the box; add your own allowance on top, as we did, and they float. The MCP Apps specification describes safeAreaInsets only as “Safe area boundaries in pixels”, and OpenAI’s docs say the composer stays overlaid in fullscreen without giving its size (both checked October 7, 2026). Our 47 px and 127 px are estimates from screenshots.
On October 6 a conversation on a phone showed our player under ChatGPT’s title bar, so we moved the header below the top inset and the player to the bottom, where our CSS also kept 96 px for a message box on top of the bottom inset. A local test with a 34 px inset looked right; ChatGPT’s 127 + 96 did not. Now the panel keeps whichever is larger, and reads ChatGPT’s window.openai.safeArea too:
// panel.js, simplified: the host's inset or 96 px, never both
const bottom = Math.max(context.safeAreaInsets?.bottom || 0, window.openai?.safeArea?.insets?.bottom || 0);
root.style.setProperty('--safe-bottom', `${bottom}px`);
root.style.setProperty('--composer-space', fullscreen ? `${Math.max(0, 96 - bottom)}px` : '0px');
/* panel.css */
.stage > .player { bottom: calc(10px + var(--safe-bottom) + var(--composer-space)); }
What changes while a fullscreen MCP App is open on a phone?
The insets. In our run the bottom edge moved when a draft grew to two lines and when the Thinking bar replaced the message box. The specification lets a host send ui/notifications/host-context-changed whenever a context field changes, with only the changed fields, for the view to merge.
Our player sits on a CSS variable and followed. The selection card is placed by script and stayed behind, so now every context change places it again. Our own notice caused a second drift: “Updated from the chat” shows above the score for 20 seconds and pushes the score area down. The card, measured from that area’s top, slid over the player, and if anything placed it again meanwhile, it floated above the player once the notice went. A ResizeObserver on the area now places the card too, once per animation frame:
this.context = {...this.context, ...params}; // host.js, simplified: merge the partial update
function hostContext(next) { applyContext(next); editor.queuePlace(); }
new ResizeObserver(() => { updateClip(); editor.queuePlace(); }).observe($('stage'));
Why does an MCP App show an old version after leaving fullscreen?
Because ChatGPT on the phone built a new panel each time it entered or left fullscreen, and gave it the conversation’s original tool result, made before any edits. The specification lets a host tear a view down at any point; a tool result only records the moment its tool ran.
Ours was half snapshot, half live: page images and notes for its own revision, plus a current-revision field read when the panel fetches it. The rebuilt panel drew revision 0, saw that the score was at revision 3 and went read-only. Its check for chat edits compared the server’s revision with that already-current field, so it never caught up. Now a panel that the host hands a result the score has moved past reads the current revision once, unless the result holds a suggested change awaiting approval, which is meant to differ.
How do you test ChatGPT’s phone behavior without a phone?
Imitate the real host in a replica, and prove each test fails without its fix. We use sunpeak, which replicates ChatGPT and Claude for Playwright tests. Its mobile ChatGPT shell (version 0.20.91) draws a title bar and message box and reports the box as a 92 px inset. The three tests we wrote for these bugs, on a 402 × 874 touch screen, add the rest. Through the shell’s sandbox frame, which is the panel’s parent as ChatGPT’s frame is, they send new insets of 127, 160, 70 and 127 px, and after an edit they reload the panel and hand it the original tool result. Another test edits the score outside the panel, as the chat does. The player must end 0–24 px above the larger of the inset and 96 px, the card 0–16 px above the player.
- Run the real hostChatGPT’s mobile site in a Simulator
- Fix one causeDeploy, refresh the app’s tools, look again
- Imitate the hostSend its messages from the replica’s frame
- Undo the fixThe new test must fail without it
| Fix undone | What failed |
|---|---|
| The inset or 96 px, never both | Message-box test: “player close to the message box” |
| Card placed on a context change | Message-box test: “card clear of the player” |
| Current version for a rebuilt panel | Rebuild test: undo stayed disabled |
| Card placed when the area resizes | Notice test: “card docked on the player” |
A run of those three took 59.9 seconds. A fourth, added the same day, checks the phone editor’s one-row header and player, and our contributor rules now require the phone tests after any change to the fullscreen layout, insets, card or result handling. The phone run also caught ChatGPT turning a selected B4 into C4, a seventh down, for “Change this note to C”; a nearest-octave rule in the tool’s description took a local test with a real model from 2 of 3 passing runs to 6 of 6.
What we haven’t verified
Our phone checks ran in the iOS Simulator and ChatGPT’s mobile website: no physical phone, no native app. Still open:
- the exact insets ChatGPT sends, and when; ours are estimates the replicas reuse;
- Safari’s expanded toolbar, which half hides the message box and the player’s lower edge;
- a page scroll of about 156 points Safari left after the keyboard closed, which a cross-origin panel can’t undo;
- an opaque band up to about 777 points when fullscreen opened while ChatGPT was still answering;
- ChatGPT’s reply sheet and its own iframe teardown, which the replicas don’t imitate;
- ChatGPT on a computer, where fullscreen is a side panel with the message box outside it; we didn’t read its insets.
A checklist for MCP Apps on phones
- Keep both insets clear; if you also keep a fallback for a message box, use the larger, never the sum.
- Merge each
host-context-changednotification, and place again anything positioned by script, once per frame. - Observe your own containers for notices that move them.
- Expect a rebuild from an old tool result; compare its version with the server’s on load.
- Give your local test host the insets measured on the real one.
- Treat the UI resource URI as a cache key, as OpenAI’s docs advise, and after publishing a new one refresh the app’s tools in ChatGPT (plugin settings, Manage app, Refresh tools; checked October 7, 2026).
- Turn each real-host finding into a replica test that fails without its fix.
Sources
- SEP-1865: MCP Apps: Interactive User Interfaces for MCP (stable, 2026-01-26) — Model Context Protocol
- Add UI to your MCP server — OpenAI
- UI guidelines — OpenAI
- Reference (the
window.openaicomponent bridge) — OpenAI - MCP Testing Framework — sunpeak
Questions and answers
Does ChatGPT count its message box in safeAreaInsets on a phone?
In our run, yes. On ChatGPT’s mobile website (iPhone 17 Simulator, October 7, 2026) the bottom inset was about 127 px and covered the message box, and the top inset, about 47 px, covered the title bar. The MCP Apps specification doesn’t say what an inset contains, so keep the reported space clear and don’t add an allowance of your own on top of it.
Does ChatGPT change the insets while a fullscreen app is open?
Yes, in our run: a two-line draft and ChatGPT’s Thinking bar both moved the bottom edge. Apply new insets whenever ui/notifications/host-context-changed arrives, merging the fields it carries into what you have, and place again anything you position with script.
Why does my MCP App show old data after switching to fullscreen on a phone?
ChatGPT on the phone built a new panel from the conversation’s original tool result each time it entered or left fullscreen, so ours showed the score as it was before later edits. When the view loads, compare the result’s version with the server’s and read the current state once if the result is behind.
Can I test ChatGPT’s phone layout without a phone?
Partly. We replay what the phone showed in sunpeak’s mobile ChatGPT shell with Playwright: moving insets, a rebuilt panel and an edit from the chat, and each test fails with its fix undone. The values ChatGPT really sends, Safari’s toolbar, the keyboard, ChatGPT’s reply sheet and the native app still need a real phone.