Agentic Research

WeChat Open Source WeMM-Embedding Deep Dive: The Multimodal Embedding Model Topping MMEB-v2, Can It Run on Your Mac?

2026/08/3141 min readBryan Chan閱讀中文原文
TopicsWeMM-EmbeddingTencentMultimodalEmbeddingRetrieval

One-sentence summary

WeMM-Embedding is a family of general-purpose multimodal embedding models open-sourced by the Tencent WeChat Vision Team (WeChat Vision Team) (2B/4B/9B, Apache 2.0). The 9B version tops the official MMEB-v2 leaderboard with a score of 80.6, while its 2B version can already run locally on a single Mac Studio at 47 ms/item.

This article breaks it down across three levels: technical report methodology (why it is strong), engineering readiness (how usable it is), and local testing (whether it can be adopted into our infrastructure).


Project Overview

ItemValue
Produced byTencent WeChat Vision Team (WeChat Vision Team)
Technical reportarXiv:2608.24053 (2026-08-25)
Open-source licenseApache 2.0 (Tencent-authored code)
Model specifications2B / 4B / 9B, based on the Qwen3.5 native multimodal backbone
Supported modalitiesText, image, video, visual documents, arbitrary interleaved multimodal (audio not supported)
Output dimensionsMatryoshka nesting: 64-4096 (depending on specification)
Leaderboard resultMMEB-v2 overall score 80.6 (9B), ranked first on the official leaderboard (as of 2026-08-24)
Production deploymentWeChat Channels, Official Accounts, Moments, e-commerce recommendations + WeChat Search, with positive gains across all 14 online A/B tests

Why This Matters

Embedding models are the foundation of retrieval, recommendation, RAG, and agent memory systems. Previously, this foundation had two fault lines:

  1. Modality fault line: Text embeddings (BGE, E5 series) and visual embeddings (CLIP family) each went their own way, so files mixing image + text + video could not enter the same vector space.
  2. Agent fault line: A new generation of agent systems needs unified retrieval over tool descriptions, GUI screenshots, and memory snippets, which traditional text embeddings cannot cover.

WeMM-Embedding tops two leaderboards at the same time, landing exactly on these two fault lines: first on MMEB-v2 (image/video/visual document), and also first in the Text and Agent groups of MMEB-v3 (tool/GUI/memory retrieval, 47 tasks). In other words, it is not just "another multimodal CLIP", but the embedding among currently public models that comes closest to being a "standard component of agent infrastructure".


Architecture Breakdown: Turning an MLLM into a Vector Machine

<embedding> token + last-token pooling

WeMM is built on the native multimodal architecture of Qwen3.5. The input can be text segments, images, and videos interleaved in any order. A dedicated <embedding> token is appended to the end of the sequence, and its last-layer hidden state is L2-normalized to produce the output vector:

S = [z_1, z_2, ..., z_N, z_emb]   →   e = h_emb / ‖h_emb₂

Under causal attention, the <embedding> token naturally attends to all preceding text and visual content, without any need for an additional pooling head or projection layer modification.

One Forward Pass, Multiple Vectors

The same causal structure allows multiple <embedding> tokens to be inserted in the middle of the sequence. The example given in the technical report is very practical: insert one after the video and another at the end of the sequence, so a single forward pass simultaneously produces a "pure video vector" and a "video + ASR text joint vector", allowing downstream recommendation systems to take what they need. This design is a very good fit for the scenario of "multiple retrieval perspectives for the same asset" (the charts in our annual report vs. the full text).

Matryoshka: One Model, Six Dimensions

After MRL (Matryoshka Representation Learning) training, taking the first d dimensions and re-normalizing them yields a lower-dimensional vector. 2B supports 64/128/256/512/1024/2048, and 9B supports up to 4096. 256 dimensions retain 98.7% of full-dimensional image/video performance, which means index storage can be cut to 1/8 with almost no performance drop.

Two-Stage Training: From "Broad Coverage" to "Fine-Grained Relevance"

Stage 1: Large-Scale Alignment with Hundreds of Millions of Pairs

A unified pair format (instruction, query, target, hard_negatives, graded_score) packs retrieval, classification, QA, and grounding into a single multi-task pipeline. Three objective functions run in parallel:

  1. InfoNCE contrastive + duplicate-aware masking: near-duplicate source/target pairs within a batch are masked out to avoid false negatives in classification tasks (ablation: -0.5 points if removed);
  2. score-gap weighted CoSENT: human-graded relevance labels are converted into ranking constraints, and the larger the gap, the higher the weight;
  3. MRL multi-dimensional loss: each batch is optimized simultaneously across all supported dimensions.

The largest impact in the ablation comes from task-consistent batching (the same batch comes from the same task and candidate space): switching to mixed sampling directly causes -3.4 points. The lesson is straightforward: the quality of negative examples in contrastive learning depends on the consistency of the candidate space within a batch, and mixed batching is a hidden killer.

Stage 2: Curated Data + Distillation + Merging

Stage 2 uses only about 1/10 of the data volume of Stage 1, but does three refined things:

  1. Semantic-ID guided resampling: use an intermediate checkpoint to encode pairs, fit three-layer RQ-KMeans to obtain Semantic IDs, and resample inversely according to codebook density to suppress repeated exposure of high-frequency semantic patterns;
  2. MLLM quality control + hard-negative expansion: a multimodal large model filters mismatched pairs and corrects factual errors in weakly supervised text; text negatives are generated by the MLLM, visual negatives are mined via retrieval with an intermediate checkpoint, and some are further scored by a reranker;
  3. Bidirectional KL embedding distillation: freeze 9B as the teacher, and perform KL distillation on the similarity distributions in both source→target and target→source directions within a batch. This is the largest source of gains for small models, and the strong performance of 2B/4B is to a large extent distilled from 9B.

9B itself has no larger teacher, so it trains multiple specialized Stage 2 variants with complementary data mixtures, and finally uses model merging to synthesize the final version.

Cumulative ablation: the stacked Stage 2 strategies total +2.2 points (curated → reranker → distillation → higher visual budget).

Another honest detail: reranker supervision is not universally effective. The report explicitly states that reranking self-retrieval results yields unstable gains on most multimodal tasks, so it is restricted to a reliable subset. This attitude of "documenting what does not work in the report" is a benchmark for technical report quality.


Benchmark Results Overview

MMEB-v2 (78 datasets)

ModelSizeAVGImageVideoVisDoc
Qwen3-VL-Embedding2B73.275.061.979.2
DME-Small†2B74.875.965.679.9
WeMM-Embedding2B77.979.670.880.7
WeMM-Embedding4B79.280.872.182.0
Qwen3-VL-Embedding8B77.880.167.182.4
WeMM-Embedding9B80.681.974.383.3

The 2B model outperforms its weight class against 8B, while the 9B model surpasses all open-source and closed-source models.

MMEB-v3 (190 tasks, including 53 text + 47 agent + 11 audio + MCMR): 2B=56.0 / 4B=58.2 / 9B=59.5, ranks first in both the Text and Agent groups. Audio is not supported and receives a score of 0, yet it still takes first place overall, showing how far ahead the other modalities are.

Cross-modal retrieval on 12 benchmarks: 2B=79.8 / 9B=81.7, going head-to-head with closed-source commercial models such as Gemini Embedding 2, Amazon Nova MME, and Voyage Multimodal 3.5.

WeChat internal 26 tasks + 14 online A/B tests: beats open-source baselines across all five task categories; in recommender systems, it is used for candidate recall, ranking features, user sequence modeling, and cross-domain understanding; gains are most pronounced for long-tail and newly published content, which is good news for cold-start scenarios (our newly archived project files).


Local Benchmark: Real Numbers on Mac Studio MPS

We fully deployed 2B on a Mac Studio (arm64, 512GB RAM) and compared it with local BGE-m3 (isolated venv, without changing any system configuration):

MetricWeMM-Embedding-2BBGE-m3
Vector dimension2048 (MRL 64-2048)1024
Text latency (after warm-up)47 ms3 ms
Image latency321 ms/image ✅❌ Text-only
Retrieval Top-3 hits (24 Chinese-English corpus items + 6 queries)6/6 (100%)6/6 (100%)
Peak MPS memory≈6.0 GB≈1.0 GB
Weights5.44 GB safetensors4.3 GB

The conclusion is clear: in text-only scenarios, BGE-m3 is about 20x faster and uses 6x less memory, so there is no reason to switch; however, multimodality is a dimension BGE-m3 can never reach, and 321 ms/image is completely acceptable for non-real-time archiving scenarios.

Testing Pitfalls (Honestly Recorded)

  1. huggingface-cli has been deprecated → switched to hf download;
  2. import torchvision triggered a SIGABRT from duplicate libomp initialization → bypassed with KMP_DUPLICATE_LIB_OK=TRUE (officially marked as an unsafe workaround, a known macOS issue);
  3. Installing torchvision also upgraded torch 2.12.0 to 2.13.0 in the venv;
  4. flash-linear-attention/causal-conv1d were not installed → linear attention used the torch fallback path; results were correct but throughput was not optimal;
  5. bf16 internal normalization error norm≈1.0026 → after re-normalizing in fp32, exactly =1.0;
  6. README recommended transformers==5.2.0; in testing, 5.9.0 already includes the qwen3_5 architecture and loads normally;
  7. The actual weights are 5.44GB, about 2x different from the 2.72GB listed in HF metadata (presumably about 2.7B parameters in bf16), which is inconsistent with the "2B" naming.

Video embedding was not tested due to missing decord, which is a known gap in this benchmark.


Implications for Agent Infrastructure: Dual-Channel Vector Memory

Based on empirical testing, our recommendation for self-built agent memory systems is a dual-channel architecture:

  1. Text channel: Continue to use BGE-m3 (3ms, 1GB), serving as the primary workload for everyday memory retrieval;
  2. Multimodal channel: Load WeMM-2B on demand (release 6GB of MPS memory after use), specifically to handle archival retrieval of images, scanned documents, and video frames; for low-dimensional deployment, use 256-dimensional MRL truncation (verified norm=1.0), cutting storage and distance computation costs by a factor of 4-8.

A deeper implication lies in the Agent group of MMEB-v3: when tool descriptions, GUI screenshots, and memory snippets need to be retrieved in the same space, agent-centric embedding will become the next-generation foundation of memory systems. WeMM currently ranks first among publicly available models in this track.

Only with a CUDA server can WeMM's flash-linear-attention fast path and 9B/4096-dimensional version reach their full potential; local Mac is "usable", while a server is "performant".


Licensing and Risks

  • Apache 2.0: Commercial use friendly; third-party components retain their original licenses, self-check required before use;
  • No audio support: MMEB-v3 audio group scores 0, omni-modal (omni) is official future work;
  • Version dependencies: transformers recommended 5.2.0 (tested 5.9.0 works), vLLM 0.27.0 / SGLang 0.5.9 serving recipes complete;
  • Naming vs actual parameter count: 2B weights 5.44GB (about 2.7B parameters), use measured values for capacity planning.

Bottom Line

WeMM-Embedding has advanced "multimodal unified vector space" from papers to the stage of WeChat production environment validation + open-source reproducibility + consumer-grade hardware runnable. For any team working on multimodal retrieval, agent memory, or content archiving, it is the open-source embedding most worth serious evaluation in 2026 Q3, bar none.

Our next step: dual-channel architecture PoC, truly putting annual report charts and scanned documents into the same retrieval space.


References: arXiv:2608.24053 · github.com/Tencent/WeMM-Embedding · huggingface.co/collections/tencent/wemm-embedding · local test report (2026-08-31, Mac Studio MPS)