Agentic Research

Read these first

This article assumes the following earlier in its learning path.

Build Your First AI Agent in Seven Steps: from environment to a minimal production pass

2026/10/0327 min readBryan Chan閱讀中文原文
TopicsAI AgentTutorialLLMTools

Most tutorials teach you "what an agent is" and stop there. This piece takes the opposite route: finish one project end to end, wiring your first API call, prompt design, function calling, the loop, RAG, and evals into a single line. The project is deliberately humble — a "paper assistant": give it a topic, it fetches abstracts, grades them, and writes a short note. But it passes through every checkpoint that actually matters.

This article adapts the integrated tutorial from awesome-agentic-ai-zh (MIT; source linked at the end) with heavily trimmed code; the full version with the cloud path and cost worksheets is linked below.

StepWhat you doOutput
Step 0Environment setupWorking Python 3.11 + Ollama
Step 1First LLM callhello_llm.py
Step 2The four-part promptStable summary output
Step 3Tool useAuto-fetching paper abstracts
Step 4Loop + reflectionA retrying agent
Step 6RAG memoryAnswers with evidence
Step 7Evals + observability + brakesA rerunnable quality check
Step 8The smallest doorA CLI interface with a safe exit

(The upstream curriculum deliberately skips step 5's ecosystem material — a from-scratch project doesn't need it yet.)

Step 0: Environment setup

Install everything later steps need, once. Ollama keeps the whole path local — API fees are 0; your electricity and hardware are still real costs.

python3 --version          # needs 3.11 or newer
python3 -m pip install "openai>=1.0,<3"
ollama pull qwen2.5:3b     # a small starter model is enough
ollama serve               # default port 11434

Create one folder for the project and commit every step to Git — step 7's "undo" depends entirely on it.

Step 1: Your first LLM call

Five core lines. The point is not pretty output but reading the token counts from usage: that is your entry point for cost and latency later.

# step1_hello_llm.py
from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
r = client.chat.completions.create(
    model="qwen2.5:3b",
    max_tokens=100,
    messages=[{"role": "user", "content": "Introduce yourself in one sentence."}],
)
print(r.choices[0].message.content)
print("output tokens:", r.usage.completion_tokens)

Step 2: Write the prompt as four parts

A vague "tidy this up" becomes four slots: goal, data, rules, output. The data changes while the rules stay — that is what makes a prompt reusable.

# step2_paper_summary.py (excerpt)
SYSTEM_PROMPT = """Goal: turn a paper abstract into three lines of research notes.
Rules: use only the data; write "not mentioned" where evidence is missing.
Output: three lines, each starting with a dash."""

def summarize(text: str) -> str:
    r = client.chat.completions.create(
        model="qwen2.5:3b",
        messages=[{"role": "system", "content": SYSTEM_PROMPT},
                  {"role": "user", "content": f"data: <input_data>{text}</input_data>"}],
    )
    return r.choices[0].message.content

Step 3: Tool use: fetching papers automatically

The model cannot reach the internet by itself. You declare a tool (its schema is the instruction card), the model replies "I want to call fetch_abstract with these arguments", and your program validates, executes, and returns the result.

# step3_tool_use.py (excerpt)
TOOLS = [{
    "type": "function",
    "function": {
        "name": "fetch_abstract",
        "description": "Fetch the abstract for an arXiv paper ID",
        "parameters": {
            "type": "object",
            "properties": {"paper_id": {"type": "string",
                                         "description": "e.g. 2210.03629"}},
            "required": ["paper_id"],
            "additionalProperties": False,
        },
    },
}]

Two iron rules: dispatch only tool names on the allowlist; treat arguments as untrusted input. The full five-step round trip matches the function calling primer.

Step 4: Add the loop and reflection

Wrapping step 3's single round trip in a loop gives you the minimal agent loop. The cap (MAX_STEPS) is not decoration — without it, a stuck model loops with your program until dawn.

# step4_loop.py (the core 13 lines)
for step in range(MAX_STEPS):
    resp = ask_model(messages, tools)
    calls = read_tool_calls(resp)
    if not calls:
        break  # no tool request → treat as done
    for call in calls:
        name, args, call_id = validate_call(call)   # allowlist + argument checks
        result = TOOL_IMPL[name](**args)
        messages.append(make_tool_result(call_id, result))
else:
    raise RuntimeError(f"Exceeded {MAX_STEPS} steps, stopping")

Start with the cheapest reflection: if the output misses its "not mentioned" markers or the line count is wrong, feed the error description back as the next user message. Change one thing at a time so you know which change worked.

Step 6: Add RAG memory

So far the model starts from zero every time. RAG gives it a notebook it can consult: chunk fetched abstracts, embed them, store them, and retrieve the most relevant chunks into the prompt before answering.

# step6_rag.py (excerpt)
collection.add(ids=[paper_id], documents=[abstract])  # store into Chroma
hits = collection.query(query_texts=[question], n_results=3)
context = "\n---\n".join(hits["documents"][0])
# put context into the prompt and require citation markers

Two honesty rules: answers must cite sources; say "I don't know" when retrieval finds nothing. RAG is not memory — persisting preferences across sessions is a separate thing; see RAG deep dive for details.

Step 7: Evals, observability, and brakes

Without this step, everything before it only "worked once in front of you". Three things:

# step7_eval.py (excerpt)
CASES = [
    {"q": "What is ReAct?", "must_contain": ["Reasoning", "Acting"]},
    {"q": "some made-up term", "must_contain": ["don't know"]},
]
for c in CASES:
    answer = agent_answer(c["q"])
    ok = any(k in answer for k in c["must_contain"])
    print(("PASS" if ok else "FAIL"), c["q"])

# log one line per run: time, steps, tokens, result
log.write(f"{ts}\t{steps}\t{tokens}\t{ok}\n")

Fixed cases must not change mid-experiment, or two scores cannot be compared; high-risk actions (email, deletion, payment) get human confirmation; every run leaves an inspectable record. This is exactly the minimal eval from the evaluation guide.

Step 8: Pick the smallest door

The final step is choosing an interface. The order: just finding data → web search/fetch; work lives in web pages → browser use; across desktop apps → computer use; running untrusted code → sandbox. Never open a bigger door than the task needs. This project needs only a CLI — one argparse entry point is a legitimate interface.

The rails: safety rules that hold across all seven steps

Whichever step you are on, five rules never change: execute only allowlisted tools; treat arguments as untrusted; give tools least privilege; ask a human before high-risk actions; set caps (turns, timeout, cost). They come from the upstream curriculum's repeated warnings and are the heart of this site's agent harness concept.

Common stuck points and fixes

  • The model won't call the tool: rerun three times with the same prompt, model, and schema before concluding anything; one failure never proves "unsupported".
  • Responses get truncated: shorten the input or lower max_tokens, and check the model's context limit.
  • Tool results don't match up: call IDs were not wired back — bind every tool result to its original request per the spec.
  • Scores swing wildly: too few trials; run the same case at least five times and look at the distribution.

Next steps

  • Map these seven steps onto this site's learning roadmap: every step has a dedicated stop to go deeper (first LLM call, function calling, RAG, evals).
  • For the full version (cloud path, cost worksheets, line-by-line notes): the upstream seven-step tutorial lives in awesome-agentic-ai-zh (MIT).
  • When you finish, take it to the roadmap's capstone: swap the five eval cases for questions from your own field.

Next on this path