LLM Basics
How do LLMs, tokens, context, and models differ?
Call an LLM and choose cloud or local models to fit the need.
📌 Learning goals
- Make your first API call with a local Ollama model, then compare it with Anthropic.
- Explain the order from pre-training and post-training to inference.
- Explain token, context window, and temperature with simple examples.
- Read input and output token counts from a response's usage field.
- Explain a model choice using input/output price, latency, and data sensitivity.
Entry conditions
Start here directly if you already know Python, Git, and APIs; otherwise go back to Stage 0 first.
🧭 Lessons on this site
Read in the suggested order; checkboxes share the same browser progress as the /learn track pages.- 01Understanding the Transformer Architecture: Attention Is All You Need
An in-depth understanding of the core components of Transformer: Self-Attention, Multi-Head Attention, Positional Encoding, and how they replaced RNNs to become the foundational architecture of modern LLMs.
7 min - 02Your First LLM API Call: Tokens, Billing, and Common Errors
Send your first LLM request with Python's openai SDK: understand the division of labor among system, user, and assistant message roles; the trade-offs of temperature, max_tokens, and streaming; token billing concepts and usage estimation; and finally, a troubleshooting guide for the errors every beginner will meet.
21 min - 03After You Press Enter: From 0.03 s to 0.85 s, What Your Question Goes Through
One real API call dissects the full chain of a question: 35 ms to establish the connection, microseconds to tokenize, 474 ms for the model to read your whole question, 380 ms generating it token by token — 854 ms in total. Includes the complete reproducible measurement method, and one counterintuitive conclusion: 96% of the time is not in the network but inside the model.
13 min - 04How AI Pricing Works: Why One Month Could Cost Ten Times More
What a token is, why long conversations get more expensive, what free tiers usually limit, how to choose between subscriptions and pay-as-you-go, and five concrete ways to keep your first month's cost near zero. No specific prices listed — this site only writes dollar amounts after manual verification.
7 min - 05May 2026 LLM API Pricing Landscape: Complete Comparison of DeepSeek, Qwen, GLM, Kimi, MiniMax, and Doubao
Cross-verified through multiple search engines including Exa, Tavily, and Brave, this compiles per-million-token pricing for DeepSeek V4, Qwen3.6-plus, GLM-5, Kimi K2.5, MiniMax M2.7, Volcano Engine Doubao Seed 2.0, and Claude/GPT, with a Sub2API cost setup guide.
12 min - 06Three-Layer Framework LLM API Configuration Summary: DS V4 Pro, Qwen 3.6, oMLX Local
A complete guide to configuring DeepSeek V4 Pro, Qwen 3.6-Plus, and oMLX local Qwen3 Coder Next across Hermes Agent, OpenClaw, and Claude Code.
9 min - 07Ollama: The Best Way to Run LLMs Locally
Ollama lets you run open source models like Llama, Mistral, and Qwen on Mac/Windows/Linux with one click. It is completely offline, free, and privacy safe.
6 min
📚 Required reading
- 1.OpenAI:模型如何開發⭐⭐⭐⭐⭐Start with how data, training, and models relate.
- 2.Google ML Crash Course:LLM 調整⭐⭐⭐⭐⭐Tell prompt engineering, fine-tuning, and distillation apart.
- 3.Anthropic Claude 模型總覽⭐⭐⭐⭐Entry point for models, context, and pricing.
🎯 Curated resources
| Resource | Who it's for | Priority | Why |
|---|---|---|---|
Official API intro Anthropic Cookbook | Need API examples | ⭐⭐⭐⭐ | Claude API notebooks for tool use, batch, and prompt cache. |
Official API intro Anthropic Courses | Want the official course | ⭐⭐⭐⭐ | Anthropic's official courses, starting with API fundamentals. |
Official API intro OpenAI Cookbook | On the OpenAI path | ⭐⭐⭐⭐ | OpenAI API, structured output, and function-calling examples. |
Chinese materials datawhalechina/happy-llm | Want the principles in Chinese | ⭐⭐⭐⭐ | Understand LLM principles and training in Chinese. |
Chinese materials datawhalechina/llm-universe | Extending toward knowledge bases | ⭐⭐⭐⭐ | Extends from API basics to knowledge bases and RAG. |
English courses Hugging Face — LLM Course | Want the ecosystem systematically | ⭐⭐⭐⭐ | Transformers, tokenizers, and the Hugging Face ecosystem. |
English courses LangChain Academy | Want a free official course | ⭐⭐⭐ | Official free course including RAG and agents. |
Local runtimes ollama/ollama | On the local Path A | ⭐⭐⭐⭐ | The local runtime used by this stage's Path A. |
Local runtimes ggml-org/llama.cpp | Want the quantization layer | ⭐⭐⭐⭐ | Understand quantization and the local inference layer. |
From scratch Karpathy — Let's build GPT from scratch | Learn by watching a build | ⭐⭐⭐⭐ | Build a GPT from scratch with PyTorch. |
From scratch rasbt/LLMs-from-scratch | Want a book-length walkthrough | ⭐⭐⭐⭐ | Work through tokenizers, attention, and training with code. |
From scratch jingyaogong/minimind | Want to train a tiny model yourself | ⭐⭐⭐ | Implement a small model from scratch; Apache-2.0. |
🛠 Hands-on practice (upstream)
Full exercises & starter codeSummaries from the upstream curriculum; full code, cost, and latency estimates live upstream.
- Exercise 1: LLM API hello world — get a response with a five-line core call and read output tokens from usage (local Ollama path costs $0).
- Exercise 2: tokens — measure input/output tokens on the same text and estimate a per-call cost.
- Exercise 3: pricing/latency — compute a cloud-call cost from measured tokens and official rates, and time the latency.
- Exercises 4–6: cross-provider comparison, error handling, local LLM — compare providers on one prompt, handle connection and truncation errors, and run a local model.
- Recommended capstone: a personal document-summary cost/quality comparator — summarize with Ollama and one cloud model, record tokens, latency, and cost, and grade quality with a fixed checklist.
✅ Self-check
- Explain what API, token, and context window each do.
- Run Exercise 1's Ollama Path A and read output tokens from usage.
- Calculate one cloud-call cost from measured input/output tokens.
- Explain why you chose local or cloud for one scenario and name one limitation.
Adapted from awesome-agentic-ai-zh (MIT, by Wenyu Chiou) v2026.09.23; links checked 2026-08-27. Stars mark learning priority (⭐⭐⭐⭐⭐ = you will get stuck without it), not popularity. MIT License · Curriculum structure last updated 2026-10-03. Content is still being filled in; lessons marked “in progress” are not live yet.