Web Fetch and Web Scraping in Agent Applications
The Role of Web Fetch in the Three Layers
| Framework | Purpose | Method |
|---|---|---|
| Hermes Agent | Background information gathering | Python requests / fetch_url |
| OpenClaw | Fetching web pages during tasks | MCP Puppeteer / fetch_url |
| Claude Code | Used indirectly through OC | Does not fetch directly |
Static Page Fetching (Lightweight)
# Python fetch (for Hermes background collection)
import requests
from bs4 import BeautifulSoup
resp = requests.get("https://example.com/article")
soup = BeautifulSoup(resp.text, "html.parser")
text = soup.get_text()
Dynamic Page Scraping (MCP Puppeteer)
For pages that require JavaScript rendering, use Puppeteer MCP:
{
"mcpServers": {
"puppeteer": {
"command": "npx",
"args": ["@anthropic-ai/mcp-server-puppeteer"]
}
}
}
Typical Scenarios
| Scenario | What to Use | Example |
|---|---|---|
| Scrape AI news summaries | Hermes + requests | Daily collection |
| Query HKEX announcements | OpenClaw + Puppeteer | In the /sdd workflow |
| Verify whether external links are broken | Script + requests | Check during build |
Determine Static or Dynamic First
Choosing the wrong method is the most common source of waste. Lightweight scraping is fast and low-cost, but it cannot retrieve content rendered by JavaScript; browser scraping can handle dynamic pages, but it is resource-intensive and slow. The recommended decision sequence is as follows.
- First scrape once using the lightweight method, and check whether the response content already contains the target data.
- If the target data is not in the raw HTML source (for example, populated by a frontend framework after loading), switch to browser rendering.
- If the page requires interaction (clicking, scrolling, logging in) before it displays the data, only then upgrade to full browser automation.
| Method | Advantages | Limitations |
|---|---|---|
| requests with parser | Fast, low overhead, easy to parallelize | Cannot retrieve rendered content |
| Browser rendering | Can handle dynamic and interactive content | Slow, resource-intensive, susceptible to layout changes |
| Official API | Stable, well-defined fields | Requires authorization, may have quotas |
As long as authorization can be obtained, an official API is always preferred over scraping. Scraping is suitable when there is no API or the API is insufficient, and it is a supplementary approach.
Compliance, Politeness, and Anti-Scraping Trade-offs
Before scraping, confirm the target site's terms of use and robots rules, and avoid putting pressure on the source.
- Set a reasonable request interval, and do not hammer the same site intensively within a short period.
- Include an identifiable User-Agent so the other party can identify and contact you.
- Respect robots rules and any scraping restrictions declared by the website.
- When encountering rate limiting or blocking, use exponential backoff retries instead of continuing to retry aggressively.
- Before large-scale scraping, assess whether authorization or a paid plan is required.
- Pay attention to personal data and copyright: take only what is needed, and do not copy and store the entire site.
Anti-scraping mechanisms are common on dynamic websites and may include CAPTCHAs, behavioral detection, and fingerprinting. The correct approach to handling these mechanisms is to switch to official channels or reduce frequency, not to bypass protections. Bypassing protections is unreliable and can easily create legal and reputational risks.
Content Cleaning and Storage
Fetched HTML often mixes navigation bars, ads, footers, and scripts; feeding it directly to a model both wastes tokens and reduces accuracy. Common practices include:
- Extract the main content block and discard boilerplate content (which can be determined by tag structure or known main content containers).
- Remove excess whitespace and duplicate paragraphs, and retain heading levels to make the structure understandable.
- Keep the source URL and scrape time for later tracing and citation.
- For content that needs to be cited, keep the original text rather than only a summary, so it can be verified later.
- Set a maximum response size to avoid fetching oversized files that drag down the pipeline.
Another key point is timeout and error handling: scraping without timeouts can stall the entire pipeline; error handling that does not distinguish between "page does not exist" and "temporary failure" makes the retry strategy meaningless.
Verification Methods and Common Errors
Verification methods
- Compare against a page known to render content, confirming that the lightweight method cannot retrieve it while the browser method can.
- Check whether the scraping result contains main content keywords, confirming that an empty shell page was not retrieved.
- Test timeout and failure paths, confirming that error messages are clear and that failures do not pass silently.
- Verify that the link checking script can correctly distinguish broken links from temporary connection failures.
- Spot-check whether saved content includes a source and timestamp.
Common errors
- Using browser rendering for every page, which wastes resources and slows down the process.
- Not setting timeouts for scraping, causing the entire process to hang when encountering slow sites.
- Ignoring robots rules and terms, or making excessive requests to the source.
- Failing to retain source and time, making it impossible to cite or verify later.
- Scraping only once and assuming the data is current, without considering caching and update frequency.
- Stuffing the entire fetched page content unchanged into the prompt, diluting the key information.
- Reporting temporary network errors as broken links during link checking.
Checklist
- Confirmed whether the target page requires JavaScript rendering.
- Prioritized assessing whether an official API can replace scraping.
- Confirmed that terms and robots rules permit scraping.
- Configured request intervals, User-Agent, and backoff retries.
- Configured timeouts and clear error classification.
- Extracted main content and removed boilerplate.
- Each result retains the source URL and fetch time.
- Tested critical paths such as rendering, timeouts, and broken links.
Related Articles
More in Tools
- PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds
- Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies
- Academic Research Skills Deep Technical Breakdown: How 45+ Agents Collaborate to Complete the Full Workflow from Literature Review to Peer Review
- AI Engineering from Scratch Deep Dive: 435 Lessons × 20 Stages