Mobile MCP: AI Controls Phones via Accessibility Trees, Not Pixels — 6.9k Stars, One API for iOS/Android
Mobile MCP is an open-source MCP server that lets AI agents automate iOS and Android devices by reading native accessibility trees instead of relying on vision models, providing 32 unified tools for testing, data extraction, and debugging across simulators, emulators, and real devices.
One-Line Positioning
Mobile MCP is a mobile automation MCP server : install it and your AI agent (Claude Code, Cursor, Codex, etc.) can directly control iOS/Android simulators, emulators, and USB-connected real devices — tap buttons, fill forms, install apps, fetch logs, retrieve crash reports — for automated testing and data extraction.
The key difference from traditional approaches: one API works across all targets , eliminating the need to write platform-specific glue like XCUITest or Espresso and removing the requirement to understand cross-platform differences.
Three Core Highlights
Capability: 32 tools covering the full pipeline. Grouped from the README: device management (6 tools — list devices, screen size, GPS, clipboard), app management (6 — install/uninstall/launch/stop), screen interaction (9 — screenshot, list elements, click/double-click/long-press/swipe, record screen), input navigation (3), logs & crashes (4), cloud real devices (4), plus batch execution. From tapping to crash logs, the entire chain is covered.
Architecture: Accessibility-first, saving tokens and preventing hallucinations. This is the highest-value design decision. The server prefers reading the native accessibility tree (the OS-maintained UI element structure for visually impaired users, containing each control's coordinates, text, type), obtaining structured data about real UI elements — not model guesses from screenshots. It falls back to "screenshot + coordinate click" only when the accessibility tree is incomplete. This saves image tokens, responds faster, and avoids the ambiguity of pure vision-based approaches.
Ecosystem: Full MCP client compatibility + cloud real devices. Works with Claude Code, Codex, Gemini, Copilot, Cursor, Cline, Goose, opencode, Windsurf, and other MCP clients. No local devices? The official Mobile Next Cloud offers on-demand cloud real-device rental, suitable for CI/CD and scale. The ecosystem also includes mobilewright (turns agent exploration into deterministic tests, "Playwright for mobile") and mobilecli (underlying universal device CLI).
Why "Reading the Tree" Beats "Looking at Pictures"
The old path: show a screenshot to a multimodal model → model returns coordinates → script clicks. Three inherent problems: image tokens are expensive, coordinates drift (resolution/DPR conversion errors), and models may hallucinate non-existent buttons.
Mobile MCP flips this order. Per the official architecture diagram:
(Official architecture diagram: MCP client connects via API layer to Mobile MCP server, which interfaces with local real devices/simulators, local filesystem, and cloud device farms)
Three key layers:
Device Abstraction Layer : iOS uses xcrun simctl (macOS built-in simulator manager) and WebDriverAgent; Android uses adb (Android Debug Bridge). Platform differences are absorbed here, yielding a unified upper API.
Interaction Layer : Prefers mobile_list_elements_on_screen to read the accessibility tree, returning a structured element list — real coordinates, real attributes, deterministic output. The model faces explicit targets like "Element 3: Login button" instead of pixel guessing. Only when needed does it call mobile_click_on_screen_at_coordinates as a coordinate fallback.
Service Layer : Beyond stdio local mode, supports Streamable HTTP mode ( --listen 3000) with Bearer Token authentication. This means the MCP server can be centrally deployed on a machine packed with devices, shared remotely by the whole team.
In one sentence: Turns "AI watches screen" into "AI queries structured data" — a reliability leap from guessing to querying.
Quick Local Start
Two steps to run, using Claude Code as example:
# 1. Register MCP server (first run auto-fetches via npx)
claude mcp add mobile-mcp -- npx -y @mobilenext/mobile-mcp@latest
# 2. Talk to the agent in plain language
# "List my connected devices, then open Settings app and take a screenshot"Prerequisites per platform:
iOS Simulator: macOS + Xcode (xcrun simctl)
Android Emulator: Android SDK + running emulator (adb)
Real Device: USB connection + iOS trust authorization / Android USB debugging enabledThree high-frequency use cases:
# Automated testing
"Open the app, walk through the login flow, save a screenshot at each step"
# Data extraction
"Read the current screen's list elements, extract titles and prices into JSON"
# Troubleshooting
"Grab this app's recent crash logs and summarize the crash cause"Codex / Gemini CLI install commands are in the README; npx -y @mobilenext/mobile-mcp@latest works universally for any MCP client.
Enterprise Team Adoption
1) Team transformation mindset. Positioned as "mobile automation infrastructure", not a test platform itself. Recommended approach: test teams hand repetitive manual scenarios (regression checklists, multi-step user journeys) to agent + mobile-mcp for draft execution, humans review results. Define boundaries clearly — it covers "functional flow works", not performance testing or pixel-level visual regression.
2) Deployment options. Two paths: individual developers use stdio mode (npx launch, each uses their own); teams use HTTP mode for centralized deployment — one Mac for iOS real devices, one Linux for Android devices, configure MOBILEMCP_AUTH with Bearer Token, whole team connects MCP clients remotely. Device resources managed centrally, no need for everyone to plug in phones.
3) CI/CD integration. Two choices in pipelines: cloud real devices (Mobile Next Cloud, on-demand rental, good for multi-device matrices) or local headless simulators (README has dedicated section). Recordings and crash logs flow directly into build artifacts. Advanced move: use official mobilewright to solidify agent exploratory tests into deterministic test cases for regular regression.
4) Team policy customization. Environment variables trim behavior: MOBILEMCP_DISABLE_TELEMETRY=1 disables telemetry (mandatory for compliance teams) MOBILEMCP_ALLOW_UNSAFE_URLS controls deep-link allowlist MOBILEMCP_LEGACY_ROBOT switches to legacy platform driver
Repo also includes skills/mobile-automation skill pack to codify team-specific testing conventions.
Real-World Applicable Scenarios
Mobile automated testing — regression checklists run by agent, humans only review results
App data extraction — read accessibility tree for structured data, more reliable than screenshot OCR
Data entry & form automation — batch form filling, multi-step user journeys
Troubleshooting aid — crash logs + device logs fetched in one command
Agent-to-agent collaboration — provide "hands" capability to upper-layer agent frameworks
Pros, Cons & Pitfalls
Pros : Accessibility-first saves tokens and prevents hallucinations; single API covers iOS/Android and simulators/real devices; 32 tools complete the chain; Apache-2.0 license; cloud real devices and mobilewright ecosystem; Chinese README friendly to domestic teams.
Cons : Audience skewed to mobile dev/test niche — irrelevant if you don't do mobile; "deterministic output" depends on app's accessibility labeling quality — apps with incomplete accessibility trees fall back to screenshot approach, raising both cost and uncertainty ; cloud real devices are a commercial service; local real devices require USB + authorization, two physical hurdles.
Three Pitfalls to Avoid :
iOS real device first connection requires tapping "Trust" on the phone; CI machines must pre-pair and authorize, don't let pipelines stall halfway at the prompt.
Android emulator runs on Linux/Windows/macOS, but iOS simulator only on macOS/Linux — Windows teams can only test Android or go cloud.
Apps with poor accessibility (games, custom-drawn UIs) see degraded results; before integration, probe target app's tree completeness with mobile_list_elements_on_screen.
Closing Thought
Browser-use fixed AI operating web pages; Mobile MCP is trying to fix AI operating phones. Same principle, mobile edition: Don't let the model guess pixels; let the model query structure. For mobile developers and test engineers, it's a key to offload repetitive manual testing; for non-mobile readers, the architectural judgment that "accessibility-first beats visual approaches" is worth remembering — it holds in many automation scenarios.
Open Source: https://github.com/mobile-next/mobile-mcp
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Architecture Digest
Focusing on Java backend development, covering application architecture from top-tier internet companies (high availability, high performance, high stability), big data, machine learning, Java architecture, and other popular fields.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
