Google Gemini Agent: Task-Centric Architecture Redefines Enterprise AI Agents
Google's Gemini Agent introduces a task-centric architecture where each task gets dedicated compute, agents have distinct identities with Workspace integration, four-layer memory persists across sessions, smart routing selects optimal models, and enterprise governance includes sandboxed execution, audit logs, and cost controls.
At Gemini at Work 2026, Google Cloud announced Gemini Agent, a platform that shifts the fundamental unit of work from the machine to the task. CEO Thomas Kurian stated: "work now starts in the prompt window."
From Machine-Centric to Task-Centric
Previous agent frameworks — AutoGPT, grokbot, muse — focused on managing a single agent instance. Google makes the task the basic unit: each task receives independent compute resources that can run for hours or days, with Gemini coordinating across tasks.
One API, Three Work Modes
Chat, autonomous task completion, and code generation share a single interface. Users can dispatch tasks, schedule recurring jobs, or trigger agents on events. The paradigm: you provide a goal, not step-by-step instructions; you delegate an outcome and receive completed work.
Collaborator Agents Have Independent Identities
Gemini dynamically creates sub-agents, each with its own identity and dedicated computer, running in parallel or sequence. Each collaborator agent owns a Workspace account ( @agents.company.com email, private calendar, private Drive) and appears in the corporate directory. Colleagues can @-mention it in Chat; it replies in document comments under its own name. Access follows least privilege: you share only what the agent should see.
Two forms exist:
Ephemeral sub-agents — created for a specific task and destroyed on completion.
Persistent coworker agents — long-lived team members with defined roles, operating across days, sessions, and responsibilities.
Full Device Coverage, Including Headless Mode
Agents run on Web, iOS, Android, Windows, Mac, CLI, Google Workspace, Microsoft 365, and Slack. They can also operate headless, without any UI binding.
Cloud 7×24 — Close the Laptop, the Task Continues
All devices share a unified memory, context, and personalization graph. No re‑briefing when switching devices. Long‑running tasks survive laptop sleep or shutdown.
Four-Layer Memory
Episodic memory — current task context, retained even over multi‑day runs.
Semantic memory — structured knowledge base built from documents, conversations, and inter‑agent collaboration.
Procedural memory — learned methods for completing tasks, including skills the agent writes itself.
Autobiographical memory — complete history of everything the agent has done.
Google describes onboarding as similar to a new hire: the agent first learns your tools, team, and preferences before executing work.
Proactive, Not Waiting for Instructions
Example: an email requesting a slide deck is recognized as delegable; one click hands it off. Inbox sorting ranks by importance, not arrival time, with explanations.
Enterprise Tool & Skill Registry
Connectors for Salesforce, ServiceNow, Jira, Git, Confluence, Microsoft 365, Slack, BigQuery, Databricks, Postgres, Snowflake, desktop files, and any MCP server inside or outside the corporate network. Teams publish custom tools and skills to a shared registry for company‑wide reuse. Skills are modular, reusable prompts; Gemini ships with a global skill library, and teams or individuals can add their own.
Model‑Agnostic Routing
Each task automatically matches the most suitable model. Supported today: Gemini Argon, Gemini Flash, Anthropic Claude Opus 5.5, Claude Sonnet 5.5; more proprietary and open models coming. Decoupling agent from model means context, skills, and data stay put when the best model changes. Simple tasks use small models for cost; complex tasks use large models for quality. PayPal routes 10 million multi‑model requests weekly this way.
Governance Like Employees, Not Scripts
Every agent has a cryptographically signed identity with least‑privilege access.
Audit logs attribute actions to the agent, not a human.
All agents execute inside an Agent Sandbox with isolated network boundaries.
All traffic — ingress, egress, inter‑agent — passes through an Agent Gateway (AI network firewall).
Policies written once (e.g., "agent must not open documents labeled Need to Know") apply to all agents.
Responsibility model: personal agent actions trace to the authorizing user; shared coworker agent liability falls on the admin and the data sharers, analogous to employee accountability.
Cost Controls with Hard Limits
Smart Routing directs each task to the most cost‑effective model. Cloud Billing lets admins set spend ceilings; exceeding them pauses the project automatically with one‑click resume. Costs tracked per project, enabling department‑level chargeback. Underlying infrastructure runs on TPU v8i, delivering 80% better price‑performance than the previous generation.
Workspace Inline — Inside Docs, Mail, Chat
Gemini embeds directly in Gmail, Drive, Docs, Slides, Sheets, Chat, Calendar — sharing the same memory, skills, and policies. Example: "Schedule a meeting next week with the regional event leads" — Gemini identifies the right people from Chat groups and prior email threads, checks calendars, sends invites, and coordinates external attendees. Cross‑app continuity: research market trends, build a financial model in Sheets, generate a presentation combining both — no re‑explanation needed at each step.
Industry‑Specific & Data‑Centric
Financial services edition integrates FactSet, LSEG, S&P Global, SEC filings; 50+ base skills with confidence scores, methodology display, full data lineage, verifiable citations.
Legal edition inherits matter‑level permissions and ethical walls from NetDocuments and iManage.
Knowledge Catalog maps business definitions ("net margin", "addressable market") once for all agents. Bloomberg Media saw a 63% lift in SQL query accuracy. Smart Storage writes context directly onto unstructured storage objects — no data movement. Borderless Lakehouse enables federated queries across Amazon S3, Azure Data Lake, Databricks, and Snowflake with zero egress fees.
Closing Perspective
Google is not building a better ChatGPT or another agent framework; it is constructing agents as enterprise infrastructure. Identity, permissions, audit, sandbox, firewall, cost control — individually unglamorous, but collectively essential for enterprises to trust agents in core workflows. "One task, one computer" makes the task the atomic unit; agents collaborate via identity and permissions rather than prompt chaining. If this paradigm scales, enterprise software work patterns will shift significantly. Google has advanced the form of the enterprise work agent.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
AI Engineering
Focused on cutting‑edge product and technology information and practical experience sharing in the AI field (large models, MLOps/LLMOps, AI application development, AI infrastructure).
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
