Agent Interfaces
How does an agent act on the real world beyond APIs?
Choose between computer use, browser use, and code sandboxes.
📌 Learning goals
- Look at a task and know whether it needs search, browser use, full computer use, or isolated execution.
- Explain the eight recurring core terms in your own words.
- Before the agent acts, map the sites it may visit, the actions it may take, and where it must ask a person.
- Finish a small exercise with no login, no downloads, and no real accounts touched.
- Read a benchmark by asking what was tested, how it was scored, and how many steps were allowed — not just the score.
Entry conditions
You can start once you understand Stage 3's model-proposes → program-executes → result-returns loop; Track A may do Exercise 1 only, Track B both. The bigger the door, the more it can touch and the higher the risk — start by choosing the smallest, most inspectable door.
🧭 Lessons on this site
Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.- 01Agent Interfaces: choosing between CLI, API, browser use, computer use, and sandboxes
Which door does an agent act through? Adapted from the MIT-licensed curriculum's Stage 8 concept chapter: eight recurring terms, a 'smallest interface first' decision table, the three tables to draw before letting an agent act (allowlist, actions, approval gates), and the four questions to ask of any browser/computer-use benchmark score. The bigger the door, the higher the risk — the first principle is to pick the smallest, most inspectable one.
5 min - 02Web Fetch and Web Scraping in Agent Applications
Configuration and scenario guide for using Web Fetch / Puppeteer / fetch_url to retrieve web content in a three-tier Agent framework.
6 min - 03Agent-Reach: Giving AI Agents a Pair of Eyes Over the Entire Internet
GitHub Trending #7 · 5.2K ⭐/week · Lets AI Agents search and read content from 10+ platforms including Twitter/Reddit/YouTube/Bilibili/Xiaohongshu, with no API fees, using pure web scraping
5 min - 04What Is Doubao Work: The Office Agent That Operates a Virtual Desktop and Plugs Deeply Into Feishu
Doubao Work is ByteDance's AI office agent: a desktop client plus remote task assignment from your phone, an agent that actually operates inside a virtual desktop instead of only replying with text, deep integration with Feishu enterprise knowledge (meeting minutes, group chats, documents), and three modes — Fast, Expert and Task. With the billing model and the audiences it fits.
7 min
📚 Required reading
- 1.Anthropic — Computer Use tool⭐⭐⭐⭐⭐Understand "the model proposes actions; the application executes them."
- 2.Anthropic — Browser Use tool⭐⭐⭐⭐⭐See how page elements and pixel fallback cooperate.
- 3.OpenAI — Computer Use guide⭐⭐⭐⭐See the GA tool and its safety boundaries.
🎯 Curated resources
| Resource | Who it's for | Priority | Why |
|---|---|---|---|
Official docs OpenAI Agents SDK — Sandbox guide | Building a mutable workspace | ⭐⭐⭐⭐ | Sandbox Agents remain beta; the old preview shape is deprecated. |
Browser use microsoft/playwright-mcp | Controllable browser automation | ⭐⭐⭐⭐⭐ | Check container, credentials, and framework first; origins, permissions, and data still need limits. |
Browser use browser-use/browser-use | Want a ready framework | ⭐⭐⭐⭐⭐ | Verify the actual stack against the README and releases. |
Computer use bytedance/UI-TARS-desktop | Want a desktop GUI agent | ⭐⭐⭐⭐ | See how screenshot-driven desktop action handles isolation and observation. |
Computer use trycua/cua | The macOS/Linux VM route | ⭐⭐⭐⭐ | Run computer use inside a VM to lower the risk of damaging the host. |
Sandboxes e2b-dev/E2B | Need a code-execution sandbox | ⭐⭐⭐⭐⭐ | A separate work room that only sees the files, network, and tools you put in. |
Sandboxes Modal — Sandboxes | Managed isolated execution | ⭐⭐⭐⭐⭐ | Cloud microVM sandboxes; failures are far less likely to hurt the host. |
Sandboxes Vercel Sandbox | Agents needing a remote code runtime | ⭐⭐⭐⭐ | Spin up an isolated, disposable runtime quickly. |
GUI parsing microsoft/OmniParser | Turning screenshots into actionable elements | ⭐⭐⭐⭐ | Read weights licenses per version: v3 uses an MIT YOLOv9 implementation; earlier Ultralytics detectors stay AGPL. |
Benchmarks xlang-ai/OSWorld | Evaluating tasks in real computer environments | ⭐⭐⭐⭐⭐ | Read the task set, metric, step budget, and harness together; 2.0 and the old task set percentages are not directly comparable. |
Benchmarks web-arena-x/webarena | Evaluating self-hosted web tasks | ⭐⭐⭐⭐ | Environment setup and the evaluator affect results. |
Security research Brave — Indirect prompt injection | Building a browser-agent threat model | ⭐⭐⭐⭐ | A research demo is not proof of every product's current state; read with Perplexity's BrowseSafe for vendor responses and defenses. |
🛠 Hands-on practice (upstream)
Full exercises & starter codeSummaries from the upstream curriculum; full code, cost, and latency estimates live upstream.
- Exercise 1: the minimal browser exercise — in an isolated browser profile, have the agent do one small task on allowlisted sites like example.com (no login, no downloads).
- Exercise 2: sandbox execution — with Python 3.10+, offline and keyless, implement a minimal executor and approval gate.
- Before letting it act, draw three tables: allowed sites (allowlist), allowed actions, and must-ask-a-person points (approval gates).
- Treat page text as untrusted input (the prompt-injection line), never as a higher-priority command.
✅ Self-check
- I pick the smallest interface first instead of throwing everything at computer use.
- I can explain the eight core terms and know browser use is more than reading the DOM.
- I isolate first, write the allowlist, set approvals, then verify results and logs.
- I finished the example.com exercise without the agent leaving its allowed scope.
- When I read an OSWorld score, I also look for the tasks, metric, step budget, and harness.
Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.