PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds
Core proposition: What is the most time-consuming part of Hong Kong stock M&A due diligence? It is not analysis, but manually copying financial figures from annual report PDFs. Can PaddleOCR automate this process?
Data: 81K+ Stars · 102 text regions · 96.5% average confidence · 83-second three-statement extraction · 0 errors in key figures
Methodology: Clone → venv → Transformers engine → PP-OCRv5 test → PP-StructureV3 KMeans Bug analysis → custom-built DBSCAN Parser → end-to-end Pipeline → packaged as an AK-OCR skill
Introduction: Annual Report Data Extraction, the Bottleneck in M&A Due Diligence
In pre-M&A due diligence for Hong Kong listed companies, we need to extract large amounts of financial data from the target company's annual report: income statement (P&L), balance sheet (BS), and cash flow statement (CF). The traditional approach is:
- Download the annual report PDF from HKEXnews (hkexnews.hk)
- Manually flip to the financial statement pages
- Copy numbers line by line into Excel
- Repeatedly cross-check to ensure accuracy
For a 200-page annual report, extracting just the three main statements takes 15-20 minutes, and it is prone to human transcription errors.
PaddleOCR, Baidu's open-source document AI engine (81K+ Stars, Apache 2.0), provides three core capabilities:
| Module | Function | Scale |
|---|---|---|
| PP-OCRv5 | Plain text recognition in 100+ languages | Runs on CPU |
| PaddleOCR-VL-1.6 | 0.9B-parameter VLM, OmniDocBench 96.33% SOTA | Multimodal document understanding |
| PP-StructureV3 | Full-element structuring of tables, formulas, seals, and charts | Complete layout analysis |
Our goal: use PaddleOCR to automate financial data extraction from Hong Kong stock annual reports.
1. Environment Setup: Transformers Engine on macOS
PaddleOCR 3.x supports two inference engines: PaddlePaddle and Transformers (HuggingFace).
macOS (Apple Silicon) does not have an official PaddlePaddle wheel, but the Transformers engine is fully supported. Just run:
git clone https://github.com/PaddlePaddle/PaddleOCR.git
cd PaddleOCR
python3 -m venv venv && source venv/bin/activate
pip install "paddleocr[all]" transformers torch torchvision
After all dependencies are installed, the environment requires only about 1.5GB of disk space.
II. PP-OCRv5 First Test: 00928 King International Investment FY2025 P&L Page
Test objective: page 90 of the FY2025 annual report of 00928.HK (King International Investment), the Consolidated Statement of Profit or Loss and Other Comprehensive Income.
from paddleocr import PaddleOCR
ocr = PaddleOCR(
use_doc_orientation_classify=False,
use_doc_unwarping=False,
use_textline_orientation=False,
engine="transformers",
lang="ch", # Chinese + English mixed
)
result = ocr.predict("page_90.png")
for res in result:
res.save_to_json("output") # Save structured results
Results
| Metric | Value |
|---|---|
| Text region detection | 102 items |
| Average confidence | 96.5% |
| Low confidence (<85%) | 5/102 (4.9%) |
| First inference time | ~30 seconds (including model download) |
| Subsequent inference time | ~27 seconds/page |
All key financial figures are accurate:
| Item | Original annual report | OCR extraction | Match |
|---|---|---|---|
| Revenue | 40,765 | 40,765 | ✅ |
| Cost of Sales | (28,033) | (28,033) | ✅ |
| Gross Profit | 12,732 | 12,732 | ✅ |
| Loss Before Tax | (44,906) | (44,906) | ✅ |
| Loss for Year | (47,454) | (47,454) | ✅ |
The 5 low-confidence items mainly come from single characters or concatenated English letters:
|(0.650), table separator\|I(0.328), should be11(note number)I,895(0.904), should be1,895
These errors can be handled entirely automatically with a regular expression correction table.
3. PP-StructureV3 KMeans Bug: Technical Root Cause Analysis
PP-OCRv5 can only extract text and cannot understand table structure (row and column relationships). PP-StructureV3 is a pipeline designed specifically for table structuring.
However, when we ran PP-StructureV3 on the Transformers engine, we encountered a fatal error:
sklearn.utils._param_validation.InvalidParameterError:
The 'n_clusters' parameter of KMeans must be an int in the range [1, inf).
Got 0 instead.
Root Cause Tracing
The error occurred at line 478 of paddlex/inference/pipelines/table_recognition/pipeline_v2.py:
# Call Chain
cells_det_results_reprocessing()
→ combine_rectangles(cells + ocr_miss_boxes, html_pred_boxes_nums)
→ KMeans(n_clusters=N) # N = html_pred_boxes_nums
Under the Transformers engine, the HTML table structure predicted by SLANeXt (table structure prediction model) had html_pred_boxes_nums = 0, meaning not a single cell was detected.
Passing N=0 to KMeans caused scikit-learn to directly throw InvalidParameterError.
Why Does This Happen?
The official PaddleOCR documentation clearly states: "Some models are currently still being supported," and the adaptation of the SLANeXt table structure model to the Transformers engine (PyTorch backend) has not yet been completed. It was originally designed for the PaddlePaddle engine. macOS does not have PaddlePaddle, so this path is temporarily unavailable.
4. Building Our Own Table Parser: DBSCAN Row and Column Clustering
Since PP-StructureV3 is unavailable, we need an alternative. The core idea is: use the bounding box coordinates returned by OCR and apply DBSCAN clustering to reconstruct row and column relationships.
Algorithm Design
Step 1: Y-axis DBSCAN clustering → "rows"
→ All text boxes with similar Y-coordinate midpoints are grouped into the same row
Step 2: Sort along the X-axis within each row → "Columns"
→ Sort by X coordinate within the same row
Step 3: Chinese-English Bilingual Merging
→ Feature of Hong Kong stock annual reports: a single cell contains two lines in Chinese and English
→ Merge non-numeric pairs with X-axis distance <120px
Step 4: Column Grid Detection
→ Global X-axis DBSCAN establishes a unified column boundary
Full source code (table_parser_v2.py):
import numpy as np
from sklearn.cluster import DBSCAN
from collections import defaultdict
def parse_financial_table(ocr_json, eps_y=14, eps_x=60):
texts = [fix_ocr_text(t) for t in ocr_json['rec_texts']]
boxes = ocr_json['rec_boxes']
# Step 1: Y-axis clustering → rows
y_centers = np.array([(b[1]+b[3])/2 for b in boxes]).reshape(-1,1)
labels = DBSCAN(eps=eps_y, min_samples=1).fit(y_centers).labels_
rows = defaultdict(list)
for i, label in enumerate(labels):
rows[label].append({
'text': texts[i], 'score': scores[i],
'x_center': (boxes[i][0]+boxes[i][2])/2,
'y_center': (boxes[i][1]+boxes[i][3])/2,
})
# Step 2: Sort rows by Y, sort items within row by X
sorted_rows = []
for label, items in rows.items():
items_sorted = sorted(items, key=lambda it: it['x_center'])
items_sorted = merge_bilingual(items_sorted, eps_x=120)
sorted_rows.append({
'row_y': np.mean([it['y_center'] for it in items_sorted]),
'cols': items_sorted,
})
sorted_rows.sort(key=lambda r: r['row_y'])
# Step 3: Global column grid
all_x = np.array([it['x_center'] for r in sorted_rows
for it in r['cols']]).reshape(-1,1)
x_labels = DBSCAN(eps=eps_x, min_samples=2).fit(all_x).labels_
# Step 4: Assign each item to its column
# ... (full code in repo)
Results
00928.HK annual report P&L page: 34 rows × 6 columns, fully matches the original table structure.
5. AK-OCR Pipeline: End-to-End Automation
Encapsulate the above components into a complete Pipeline:
PDF → PyMuPDF page splitting (200dpi) → PP-OCRv5 → Table Parser → JSON/MD output
End-to-End Test Results (00928 18M FY2025 Annual Report)
| Stage | Time |
|---|---|
| Phase 1: PDF → Images (3 pages) | 0.8s |
| Phase 2: PP-OCRv5 Text Recognition | 123.2s |
| Phase 3: Table Parser Structuring | ~1s |
| Phase 4: JSON + MD Output | <1s |
| Total | ~125s |
| Statement | Text Regions | Average Confidence | Structured Output |
|---|---|---|---|
| P&L (p.91) | 102 | 96.5% | 34 rows × 6 columns |
| BS (p.93) | 100 | 98.3% | 32 rows × 6 columns |
| CF (p.96) | 114 | 96.5% | 39 rows × 5 columns |
6. OCR Correction Table: Automatic Correction of Common Errors
Common misrecognition patterns of PP-OCRv5 for mixed Chinese and English documents:
OCR_FIXES = [
(r'\b3I\b', '31'), # "3I March" → "31 March"
(r'\bI,(\d)', r'1,\1'), # "I,895" → "1,895"
(r'\(2I\)', '(21)'), # "(2I)" → "(21)"
(r'\b\|I\b', '11'), # "|I" → "11"
(r'15,1\|4', '15,114'), # "15,1|4" → "15,114"
(r'diferencesarising', 'differences arising'),
(r'subsequentl y\b', 'subsequently'),
]
These correction tables can be continuously expanded as new annual report formats emerge.
7. Conclusion and Next Steps
Core Value
| Metric | Manual Extraction | AK-OCR Pipeline |
|---|---|---|
| P&L extraction from one annual report | ~15 min | ~30 sec |
| Accuracy | Risk of human error | 96.5%+ |
| Reusability | Redone each time | One-click batch |
| Cost | Manual labor | 0 (local CPU) |
Known Limitations
- PP-StructureV3 is unavailable: must wait for an official fix for Transformers engine compatibility, or switch to the Docker PaddlePaddle version
- OCR character misrecognition: digits
1↔I, touching characters; can be continuously improved with a correction table - Page number offset: annual report page numbers ≠ PDF index (difference of ~1-2 pages); automatic TOC scanning is built in
Source Code Locations
- Pipeline:
~/workspace/PaddleOCR/ak_ocr_pipeline.py - Table Parser:
~/workspace/PaddleOCR/table_parser_v2.py - Skill documentation:
~/workspace/skills/ak-ocr/SKILL.md
This article is based on real-world test data from an actual Hong Kong listed company annual report (00928.HK Emperor International Investment FY2025). All OCR results can be reproduced locally.
More in Tools
- Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies
- Academic Research Skills Deep Technical Breakdown: How 45+ Agents Collaborate to Complete the Full Workflow from Literature Review to Peer Review
- AI Engineering from Scratch Deep Dive: 435 Lessons × 20 Stages
- CodeGraph Deep Technical Breakdown: How to Save AI Coding Agents 35% in Costs and Cut Tool Calls by 70%