Financial Data API Research: Injecting Market Data Capabilities into Agent Frameworks
Research Objectives
Integrate structured financial data APIs for OpenClaw's slash commands such as /sdd (Hong Kong stock due diligence), /finance (financial analysis), and /reip (investment proposal), replacing the current manual data collection process.
Candidate APIs
Hong Kong Stock Data
| API | Data Type | Cost | Notes |
|---|---|---|---|
| HKEX Official | Announcements, annual report PDFs | Free | Requires PyMuPDF extraction |
| AASTOCKS | Real-time quotes, financial reports | Free (unofficial) | No official API |
| Bloomberg API | Full-market data | Very expensive | Institutional use |
| Morningstar | Structured financial reports | Paid | Already in use |
US Stock Data
| API | Data Type | Cost |
|---|---|---|
| Yahoo Finance | Quotes, financial reports, historical data | Free (unofficial) |
| Alpha Vantage | Quotes, technical indicators | Free, 25 requests/day |
| Polygon.io | Full-market real-time data | Paid |
| Financial Modeling Prep | Financial reports, valuation | Available on free tier |
Macro/Industry
| API | Data Type | Cost |
|---|---|---|
| FRED (Federal Reserve) | Interest rates, GDP, employment | Free |
| World Bank | Global economic indicators | Free |
| Wind | China macro/industry | Paid |
Integration Strategy
OpenClaw (/sdd 0515.HK)
→ Step 1: Yahoo Finance / Morningstar → Financial Data
→ Step 2: HKEX Official → Shareholding Structure PDF
→ Step 3: AASTOCKS scrape → Latest Announcements
→ Step 4: FRED / World Bank → Macroeconomic background
→ Consolidate → Generate research report
Priorities
| Priority | API | Purpose |
|---|---|---|
| 🔴 Immediate | Yahoo Finance | Basic data for US and Hong Kong stocks |
| 🔴 Immediate | Alpha Vantage | Fallback + technical indicators |
| 🟡 Next week | FRED | Macroeconomic background |
| 🟡 Next week | Financial Modeling Prep | Structured financial statements |
Technical Considerations and Trade-offs in the Ingestion Layer
Bringing multiple sources into the same Agent is usually difficult not because of calling APIs, but because the returned data must be comparable, traceable, and retryable. It is advisable to clarify the following areas before starting work.
Authorization and Compliance: Official APIs have clear terms and quotas, while unofficial sources (such as web scraping) carry the risk of fields changing at any time, being rate-limited, or being blocked. As a trade-off, when authorization can be obtained, prefer official interfaces; use scraping only as a supplement, and set reasonable request rates, a custom User-Agent, and retries with backoff on failures to avoid putting pressure on sources.
Caching Strategy: The update frequency of different data varies greatly. Quotes are second-level, so cache duration should be short; financial reports and annual reports are updated quarterly or annually, so cache duration can be long; macroeconomic indicators are mostly released monthly or quarterly. Setting TTL (time to live) in tiers saves quota and keeps repeated queries consistent. The key is that the cache must include source and timestamp; otherwise, stale numbers can appear in reports without being noticed.
Field Normalization: Different sources use inconsistent names and definitions for the same line item, for example, "operating revenue" and "Revenue," consolidated versus parent-company basis, and thousands versus millions as units. It is advisable to place a mapping layer between the ingestion layer and the business layer, unifying external fields into internal standard names. Business logic should depend only on internal standards, so the downstream does not need to be changed when sources are replaced.
Time, Currency, and Trading Calendars: Cross-market data involves time zones, trading calendars, and currency conversion. Hong Kong and US stock markets have different market holidays, and quote timestamps must specify the time zone; currencies from different markets must first be converted to the reporting base currency before comparison. Historical prices must also account for adjustments for ex-rights, ex-dividends, and stock splits; otherwise, long-term trends will be distorted.
PDF Extraction: Annual reports and announcements are PDFs. The extraction challenges lie in tables spanning pages, line breaks within fields, and scanned documents. You can first use PyMuPDF to extract the text layer, and consider OCR for scanned documents; extracted numbers must retain page numbers and table titles so they can be verified later.
Data Quality and Consistency Checks
The biggest risk in multi-source aggregation is that things "look right but actually use different definitions." The following checks should be automated steps rather than manual spot checks.
| Check | Approach | Purpose |
|---|---|---|
| Cross-validation | Compare two sources for the same metric | Detect single-source errors |
| Unit consistency | Explicitly label as thousands or millions | Avoid order-of-magnitude errors |
| Period alignment | Confirm whether the fiscal year is a calendar year or a custom fiscal year | Avoid period mismatch |
| Adjusted prices | Apply ex-rights and ex-dividend adjustments to historical prices | Ensure trend comparability |
| Missing-value marking | Clearly mark missing values instead of filling them with zero | Avoid treating missing data as fact |
| Source traceability | Attach source and retrieval time to each data point | Support retrospective tracing |
The tolerance for cross-validation should be set reasonably: statistical definitions from different sources may inherently differ, and rigidly requiring exact equality will produce many false positives. It is recommended to set a tolerance range, alert only when it is exceeded, and manually confirm the cause of the difference.
Practices for Integrating into an Agent Framework
Whether financial data can be used reliably by an Agent depends on tool interface design, not on the model itself.
Single-responsibility tools: Each tool should do only one thing. For example, "get quote," "get financial statements," and "get macro series" should be separate; do not create one universal query tool. The benefit of clear responsibilities is that the model can easily choose the correct tool, and debugging can quickly locate issues.
Explicit input and output structures: Tool parameters should restrict types and value ranges (for example, market code format, date format, period length), and returns should have fixed fields and include source and timestamp. A stable return format prevents downstream report generation from crashing due to missing fields.
Recoverable errors: Rate limiting, timeouts, and missing fields should return feasible alternatives, such as suggesting "use a backup source" or "shorten the query period," giving the Agent a chance to adjust on its own instead of immediately terminating the entire flow.
Common mistakes
- Hard-coding multiple sources into business logic, so changing sources later requires edits in multiple places.
- No caching, so repeated queries exhaust quota.
- Writing retrieved numbers directly into reports without cross-validation and source traceability.
- Ignoring unit and definition differences, causing order-of-magnitude or period errors.
- Unstable tool return formats, forcing the Agent to guess field meanings.
- Not handling market closures and trading suspensions; quote queries return null but are treated as zero.
- Writing secrets into configuration files or logs, causing leakage risk.
Verification Approach and Implementation Checklist
Verification Approach
First, create a set of test cases with known answers, such as the annual revenue of a specified company and the closing price on a specified date, and use these cases to verify the accuracy and field correctness of each source. Then test exception paths: network outage, rate limiting, timeout, missing fields, and market holidays, to confirm that the system reports understandable errors instead of failing silently. For caching, verify both cache hits and expiration to ensure that stale data is not returned. Finally, perform end-to-end testing, using a real slash command to walk through the entire flow and confirm that the chain from data to report is complete.
Implementation Checklist
- Licensing terms and quota limits have been confirmed for each source.
- Tiered caching has been implemented, and the cache retains the source and timestamp.
- A field mapping layer has been established, and business logic depends only on internal standard names.
- Time zones, trading calendars, and currency conversion have been handled.
- Historical prices have been adjusted for corporate actions.
- Cross-validation and tolerance alerts have been configured.
- Every data point can be traced back to its source and retrieval time.
- Exception paths such as rate limiting, timeout, missing values, and market closures have been tested.
- Secrets are managed via environment variables or protected storage and are not written to configuration files or logs.
- An end-to-end verification has been completed using real commands.
Related Articles
More in Tools
- PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds
- Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies
- Academic Research Skills Deep Technical Breakdown: How 45+ Agents Collaborate to Complete the Full Workflow from Literature Review to Peer Review
- AI Engineering from Scratch Deep Dive: 435 Lessons × 20 Stages