Agentic Research

Financial Data API Research: Injecting Market Data Capabilities into Agent Frameworks

2026/05/1031 min readBryan Chan閱讀中文原文
TopicsAPIOpenClaw

Research Objectives

Integrate structured financial data APIs for OpenClaw's slash commands such as /sdd (Hong Kong stock due diligence), /finance (financial analysis), and /reip (investment proposal), replacing the current manual data collection process.


Candidate APIs

Hong Kong Stock Data

APIData TypeCostNotes
HKEX OfficialAnnouncements, annual report PDFsFreeRequires PyMuPDF extraction
AASTOCKSReal-time quotes, financial reportsFree (unofficial)No official API
Bloomberg APIFull-market dataVery expensiveInstitutional use
MorningstarStructured financial reportsPaidAlready in use

US Stock Data

APIData TypeCost
Yahoo FinanceQuotes, financial reports, historical dataFree (unofficial)
Alpha VantageQuotes, technical indicatorsFree, 25 requests/day
Polygon.ioFull-market real-time dataPaid
Financial Modeling PrepFinancial reports, valuationAvailable on free tier

Macro/Industry

APIData TypeCost
FRED (Federal Reserve)Interest rates, GDP, employmentFree
World BankGlobal economic indicatorsFree
WindChina macro/industryPaid

Integration Strategy

OpenClaw (/sdd 0515.HK)
  → Step 1: Yahoo Finance / Morningstar → Financial Data
  → Step 2: HKEX Official → Shareholding Structure PDF
  → Step 3: AASTOCKS scrape → Latest Announcements
→ Step 4: FRED / World Bank → Macroeconomic background
  → Consolidate → Generate research report

Priorities

PriorityAPIPurpose
🔴 ImmediateYahoo FinanceBasic data for US and Hong Kong stocks
🔴 ImmediateAlpha VantageFallback + technical indicators
🟡 Next weekFREDMacroeconomic background
🟡 Next weekFinancial Modeling PrepStructured financial statements

Technical Considerations and Trade-offs in the Ingestion Layer

Bringing multiple sources into the same Agent is usually difficult not because of calling APIs, but because the returned data must be comparable, traceable, and retryable. It is advisable to clarify the following areas before starting work.

Authorization and Compliance: Official APIs have clear terms and quotas, while unofficial sources (such as web scraping) carry the risk of fields changing at any time, being rate-limited, or being blocked. As a trade-off, when authorization can be obtained, prefer official interfaces; use scraping only as a supplement, and set reasonable request rates, a custom User-Agent, and retries with backoff on failures to avoid putting pressure on sources.

Caching Strategy: The update frequency of different data varies greatly. Quotes are second-level, so cache duration should be short; financial reports and annual reports are updated quarterly or annually, so cache duration can be long; macroeconomic indicators are mostly released monthly or quarterly. Setting TTL (time to live) in tiers saves quota and keeps repeated queries consistent. The key is that the cache must include source and timestamp; otherwise, stale numbers can appear in reports without being noticed.

Field Normalization: Different sources use inconsistent names and definitions for the same line item, for example, "operating revenue" and "Revenue," consolidated versus parent-company basis, and thousands versus millions as units. It is advisable to place a mapping layer between the ingestion layer and the business layer, unifying external fields into internal standard names. Business logic should depend only on internal standards, so the downstream does not need to be changed when sources are replaced.

Time, Currency, and Trading Calendars: Cross-market data involves time zones, trading calendars, and currency conversion. Hong Kong and US stock markets have different market holidays, and quote timestamps must specify the time zone; currencies from different markets must first be converted to the reporting base currency before comparison. Historical prices must also account for adjustments for ex-rights, ex-dividends, and stock splits; otherwise, long-term trends will be distorted.

PDF Extraction: Annual reports and announcements are PDFs. The extraction challenges lie in tables spanning pages, line breaks within fields, and scanned documents. You can first use PyMuPDF to extract the text layer, and consider OCR for scanned documents; extracted numbers must retain page numbers and table titles so they can be verified later.

Data Quality and Consistency Checks

The biggest risk in multi-source aggregation is that things "look right but actually use different definitions." The following checks should be automated steps rather than manual spot checks.

CheckApproachPurpose
Cross-validationCompare two sources for the same metricDetect single-source errors
Unit consistencyExplicitly label as thousands or millionsAvoid order-of-magnitude errors
Period alignmentConfirm whether the fiscal year is a calendar year or a custom fiscal yearAvoid period mismatch
Adjusted pricesApply ex-rights and ex-dividend adjustments to historical pricesEnsure trend comparability
Missing-value markingClearly mark missing values instead of filling them with zeroAvoid treating missing data as fact
Source traceabilityAttach source and retrieval time to each data pointSupport retrospective tracing

The tolerance for cross-validation should be set reasonably: statistical definitions from different sources may inherently differ, and rigidly requiring exact equality will produce many false positives. It is recommended to set a tolerance range, alert only when it is exceeded, and manually confirm the cause of the difference.


Practices for Integrating into an Agent Framework

Whether financial data can be used reliably by an Agent depends on tool interface design, not on the model itself.

Single-responsibility tools: Each tool should do only one thing. For example, "get quote," "get financial statements," and "get macro series" should be separate; do not create one universal query tool. The benefit of clear responsibilities is that the model can easily choose the correct tool, and debugging can quickly locate issues.

Explicit input and output structures: Tool parameters should restrict types and value ranges (for example, market code format, date format, period length), and returns should have fixed fields and include source and timestamp. A stable return format prevents downstream report generation from crashing due to missing fields.

Recoverable errors: Rate limiting, timeouts, and missing fields should return feasible alternatives, such as suggesting "use a backup source" or "shorten the query period," giving the Agent a chance to adjust on its own instead of immediately terminating the entire flow.

Common mistakes

  • Hard-coding multiple sources into business logic, so changing sources later requires edits in multiple places.
  • No caching, so repeated queries exhaust quota.
  • Writing retrieved numbers directly into reports without cross-validation and source traceability.
  • Ignoring unit and definition differences, causing order-of-magnitude or period errors.
  • Unstable tool return formats, forcing the Agent to guess field meanings.
  • Not handling market closures and trading suspensions; quote queries return null but are treated as zero.
  • Writing secrets into configuration files or logs, causing leakage risk.

Verification Approach and Implementation Checklist

Verification Approach

First, create a set of test cases with known answers, such as the annual revenue of a specified company and the closing price on a specified date, and use these cases to verify the accuracy and field correctness of each source. Then test exception paths: network outage, rate limiting, timeout, missing fields, and market holidays, to confirm that the system reports understandable errors instead of failing silently. For caching, verify both cache hits and expiration to ensure that stale data is not returned. Finally, perform end-to-end testing, using a real slash command to walk through the entire flow and confirm that the chain from data to report is complete.

Implementation Checklist

  • Licensing terms and quota limits have been confirmed for each source.
  • Tiered caching has been implemented, and the cache retains the source and timestamp.
  • A field mapping layer has been established, and business logic depends only on internal standard names.
  • Time zones, trading calendars, and currency conversion have been handled.
  • Historical prices have been adjusted for corporate actions.
  • Cross-validation and tolerance alerts have been configured.
  • Every data point can be traced back to its source and retrieval time.
  • Exception paths such as rate limiting, timeout, missing values, and market closures have been tested.
  • Secrets are managed via environment variables or protected storage and are not written to configuration files or logs.
  • An end-to-end verification has been completed using real commands.

Related Articles