KnowledgeWala Workshop · Agentic AI · Day 3

Build a research agent you can trust

Gemini thinks, Tavily searches, a calculator does the maths, and the agent lists its sources. We turn a Jupyter notebook into a small project you can run, test and extend on your own laptop, then look at real results to see where agents still go wrong.

Level: beginner to intermediate Time: 15 min read · 45 min hands-on Stack: Python 3.10+, LangChain 1.x, Gemini, Tavily Cost: free tiers

In this workshop

  1. 1What changed since Day 2 Everyone
  2. 2How the agent works Everyone
  3. 3What our test runs showed Everyone
  4. 4From notebook to project Developers
  5. 5Build and run it: 8 steps Developers
  6. 6Problems and fixes Developers
  7. 710 POC ideas to build next Everyone
  8. 8Practice and quiz Everyone
Everyone

1. What changed since Day 2

In Day 2 our agent searched the web with DuckDuckGo. It worked, but it sometimes failed with network errors and returned short, messy snippets. Today we switch to Tavily, a search service built for AI agents. It returns cleaner text, a relevance score and the source URL for each result. It needs a free API key (1,000 searches a month, no card).

We also add two small tools, a calculator and today's date, plus written rules that tell the agent to cite its sources. Finally, we move the code out of one long notebook into a few short files, so it is easier to read, test and change.

PieceOlder LangChain tutorialsThis workshop
ModelOpenAI ChatOpenAIGoogle Gemini ChatGoogleGenerativeAI (free tier)
SearchSerper / DuckDuckGoTavily TavilySearch
Loading toolsload_tools()A plain Python list: [search, calculator, today]
AgentReAct PromptTemplate + AgentExecutorcreate_agent() with native tool calling
Runagent_executor.run()agent.invoke({"messages": [...]})
Code layoutOne notebookNotebook to learn, small project to run and maintain
EveryoneDevelopers

2. How the agent works

  1. 1
    YouAsk a question"What is the latest iPhone model and its price in USD?"python main.py "..."
  2. 2
    LangChainAgent adds the rules and the tool listThe question goes to Gemini together with the system prompt (search for current facts, use the calculator, cite sources) and descriptions of three tools.create_agent(model, tools, system_prompt)
  3. 3
    Gemini 3.5 Flash Lite"I need web search"Gemini recognises a question about current prices and asks for a search.tool_calls → tavily_search(query=...)
  4. 4
    ToolTavily searches the webReturns up to 5 results, each with a title, URL, text and relevance score.TavilySearch(max_results=5)
  5. 5
    ToolCalculator, if numbers are involvedExact arithmetic instead of the model's mental maths.calculator("21500 + 8400 + 3200")
  6. 6
    GeminiReads the results, decides if it needs moreNot enough? It searches again. Enough? It writes the answer.loop until no more tool_calls
  7. ✓
    YouAnswer with sourcesA plain-text answer ending in a list of the URLs it used.text_of(response["messages"][-1].content)
Gemini thinkingTool doing workLangChain / you
Everyone

3. What our test runs showed

We ran the notebook version (Gemini + Tavily, no system prompt and no calculator) on two questions in October 2026. Both runs finished without errors. One answer was right, and one contradicted its own search results.

Test A: "What is the latest iPhone model and its price in USD?"

ModelWhat the search results saidWhat the agent answered
iPhone 18 Profrom $1,199 (NewsNation, CNET)from $1,439
iPhone 18 Pro Maxfrom $1,299"$1,299 to $1,399"
iPhone Duo (foldable)from $1,999, on sale 23 Octaround $1,999

The agent searched, found the right numbers, and still gave a wrong price. Searching does not guarantee correctness. The model can still blend the results with its own guesses, which is called hallucination. In a shopping or finance tool, a $240 error like this matters.

Test B: healthcare cost check

"List the top 3 Medicare providers. A 50-day joint replacement episode costs hospital $21,500, post-acute care $8,400 and readmission $3,200. The target price is $31,000. Are we above or below target?"

What went right

Total $33,100, so $2,100 above target. The maths is correct. It also noted that "Medicare providers" was ambiguous and answered with the largest Medicare Advantage insurers (UnitedHealthcare, Humana, CVS Health/Aetna).

What to improve

The model did the maths in its head, which happened to be right this time. The output was full of LaTeX code such as \mathbf{\$33,100}, and it gave no source links for the market-share figures.

How the project fixes these

Problem we sawFixWhere in the code
Price contradicted the sourcesA rule: use only facts that appear in the search results, and say when sources disagree. List the source URLs.agent.py SYSTEM_PROMPT rules 2 and 4
Mental mathsA calculator tool, plus a rule to use it for every calculationtools.py calculator
LaTeX in the outputA rule: plain text and simple bullets onlyagent.py rule 5
Hard to see what happened--trace prints every tool call and resultoutput.py print_trace

Rules reduce these mistakes but do not remove them. For anything important, compare the answer with the sources in the trace, or keep a person in the loop.

Developers

4. From notebook to project

Notebooks are great for learning, but they get messy. Our Day 3 notebook hit NameError: name 'llm' is not defined because a cell ran before the cell that created llm. A project with small files fixes that: each file has one job, settings live in one place, and you can run it with one command and test it without API keys.

Project layout
tavily_agent/
├── .env.example     ← copy to .env, add keys
├── .gitignore       ← keeps .env out of Git
├── requirements.txt ← pinned versions
├── config.py        ← model, limits, key check
├── tools.py         ← search, calculator, today
├── agent.py         ← Gemini + tools + rules
├── output.py        ← clean text, trace
├── main.py          ← command line
├── README.md
└── tests/
    └── test_tools.py
UseWhen
NotebookLearning, trying ideas, demos in a workshop
ProjectRunning again next month, sharing on GitHub, adding tools, interviews and POCs
Developers

5. Build and run it: 8 steps

  1. Create the environment and install

    PowerShell
    # Windows PowerShell (macOS/Linux: python3, source .venv/bin/activate, cp)
    cd tavily_agent
    python -m venv .venv
    .venv\Scripts\activate
    pip install -r requirements.txt
    copy .env.example .env      # open .env and paste your two keys
    requirements.txt
    # Versions tested in the workshop notebook (Oct 2026). Pin them so the project keeps working.
    langchain==1.4.3
    langchain-core==1.6.7
    langgraph==1.2.14
    langchain-google-genai==4.4.0
    langchain-tavily==0.2.18
    python-dotenv==1.2.4
    pytest>=8

    Pinned versions are the ones the workshop notebook ran with. Pinning means the project still works in three months, even after the libraries change.

  2. Add your two free keys

    .env.example → copy to .env
    # Copy this file to ".env" and paste your real keys. Never commit .env to GitHub.
    GOOGLE_API_KEY=your_google_api_key        # https://aistudio.google.com/apikey
    TAVILY_API_KEY=your_tavily_api_key        # https://app.tavily.com  (free: 1,000 credits/month)
    
    # Optional settings
    GEMINI_MODEL=gemini-3.5-flash-lite
    SEARCH_MAX_RESULTS=5

    Get the Gemini key from Google AI Studio and the Tavily key from app.tavily.com. .gitignore already stops .env from being committed.

  3. Keep all settings in one file

    config.py
    """All settings in one place. Change the model here or in .env, never inside the code."""
    import os
    from dataclasses import dataclass
    
    from dotenv import load_dotenv
    
    load_dotenv()
    
    REQUIRED_KEYS = ("GOOGLE_API_KEY", "TAVILY_API_KEY")
    
    
    @dataclass(frozen=True)
    class Settings:
        model: str = os.getenv("GEMINI_MODEL", "gemini-3.5-flash-lite")
        search_max_results: int = int(os.getenv("SEARCH_MAX_RESULTS", "5"))
        timeout_seconds: int = 60
        max_retries: int = 3
    
    
    def check_keys() -> None:
        """Stop early with a clear message instead of a long stack trace later."""
        missing = [key for key in REQUIRED_KEYS if not os.getenv(key)]
        if missing:
            raise SystemExit(
                f"Missing {', '.join(missing)}. Copy .env.example to .env and add your keys."
            )

    When a model name stops working (Day 2 hit two model-name errors), you change one line in .env, not code in five places. check_keys() stops early with a clear message.

  4. Test search on its own first

    test_search.py (optional check)
    from dotenv import load_dotenv
    from langchain_tavily import TavilySearch
    
    load_dotenv()
    search = TavilySearch(max_results=5)
    
    result = search.invoke({"query": "latest iPhone model and price in USD"})
    for item in result["results"]:
        print(f"{item['score']:.2f}  {item['title']}\n      {item['url']}")

    Results from our run:

    Output
    0.72  Apple - Explore the latest iPhone lineup, with prices...
          https://www.facebook.com/apple/posts/...
    0.64  Apple raises iPhone prices: Here's what each model costs now
          https://www.newsnationnow.com/business/your-money/apple-raises-iphone-prices
    0.48  Best iPhone in 2026: Apple's Phones Cost More, Here Are the Ones to Buy - CNET
          https://www.cnet.com/tech/mobile/best-iphone
    0.44  Apple's iPhone Pricing is a MESS
          https://www.youtube.com/watch?v=hSfpZKA6Pzg
    0.44  Which iPhone Should You Buy? - Consumer Reports
          https://www.consumerreports.org/electronics-computers/cell-phones/which-iphone-should-you-buy-a7028071026

    Testing a tool alone before adding it to an agent tells you whether a problem is in the search or in the model. Notice that the top-scored result was a social media post, so a high score doesn't always mean the best source.

  5. Write the tools

    tools.py
    """Tools the agent may call. Each docstring is what Gemini reads to decide when to use it."""
    import ast
    import operator
    from datetime import date
    
    from langchain.tools import tool
    from langchain_tavily import TavilySearch
    
    _OPERATORS = {
        ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,
        ast.Div: operator.truediv, ast.Pow: operator.pow, ast.Mod: operator.mod,
        ast.USub: operator.neg, ast.UAdd: operator.pos,
    }
    
    
    def _evaluate(node: ast.AST) -> float:
        if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):
            return node.value
        if isinstance(node, ast.BinOp) and type(node.op) in _OPERATORS:
            return _OPERATORS[type(node.op)](_evaluate(node.left), _evaluate(node.right))
        if isinstance(node, ast.UnaryOp) and type(node.op) in _OPERATORS:
            return _OPERATORS[type(node.op)](_evaluate(node.operand))
        raise ValueError("Only numbers and + - * / % ** ( ) are allowed")
    
    
    def safe_calculate(expression: str) -> float:
        """Evaluate arithmetic without eval(), so the model cannot run arbitrary code."""
        return _evaluate(ast.parse(expression.replace(",", ""), mode="eval").body)
    
    
    @tool
    def calculator(expression: str) -> str:
        """Do exact arithmetic, e.g. '21500 + 8400 + 3200 - 31000'. Use this for every calculation."""
        try:
            return f"{expression} = {safe_calculate(expression):,.2f}"
        except (ValueError, SyntaxError, ZeroDivisionError) as err:
            return f"Could not calculate '{expression}': {err}"
    
    
    @tool
    def today() -> str:
        """Return today's date. Use it when the question depends on 'latest', 'current' or 'this year'."""
        return date.today().isoformat()
    
    
    def build_tools(max_results: int) -> list:
        web_search = TavilySearch(max_results=max_results)   # reads TAVILY_API_KEY from the environment
        return [web_search, calculator, today]

    The calculator parses the expression with Python's ast module and allows only numbers and operators. Never pass model output to eval(): a model, or a web page it read, could make it run any code on your laptop.

  6. Build the agent with rules

    agent.py
    """Builds the agent: Gemini (the brain) + tools (the hands) + rules (the system prompt)."""
    from langchain.agents import create_agent
    from langchain_google_genai import ChatGoogleGenerativeAI
    
    from config import Settings
    from tools import build_tools
    
    SYSTEM_PROMPT = """You are a careful research assistant.
    Rules:
    1. For anything that changes over time (products, prices, rankings, news, laws), use web search. Do not answer from memory.
    2. Use only facts that appear in the search results. If sources disagree or a fact is missing, say so.
    3. Use the calculator tool for every calculation, even simple ones.
    4. End with a 'Sources:' list of the URLs you used.
    5. Write plain text with simple bullet points. Do not use LaTeX or math formatting."""
    
    
    def build_agent(settings: Settings | None = None):
        settings = settings or Settings()
        llm = ChatGoogleGenerativeAI(
            model=settings.model,
            timeout=settings.timeout_seconds,
            max_retries=settings.max_retries,
        )
        return create_agent(
            model=llm,
            tools=build_tools(settings.search_max_results),
            system_prompt=SYSTEM_PROMPT,
        )

    The system prompt turns what we learned in section 3 into instructions. Each rule fixes one problem we actually saw.

  7. Run it from the terminal

    output.py
    """Helpers to turn the agent's raw response into readable text."""
    
    
    def text_of(content) -> str:
        """Gemini 3.x returns a list of content blocks; older models return a string."""
        if isinstance(content, list):
            return "".join(part.get("text", "") for part in content if isinstance(part, dict))
        return str(content)
    
    
    def print_trace(messages) -> None:
        """Show every step of the agent loop: question, tool calls, tool results, answer."""
        for msg in messages:
            kind = type(msg).__name__
            if kind == "HumanMessage":
                print(f"\n[You] {text_of(msg.content).strip()[:200]}")
            elif kind == "AIMessage" and msg.tool_calls:
                for call in msg.tool_calls:
                    print(f"[Gemini -> tool] {call['name']}({call['args']})")
            elif kind == "ToolMessage":
                print(f"[Tool result]    {text_of(msg.content)[:160].strip()}...")
            elif kind == "AIMessage":
                print("[Gemini] final answer written")
    main.py
    """Run the agent from the terminal.
    
        python main.py "What is the latest iPhone model and its price in USD?"
        python main.py --trace "Is a $33,100 episode above a $31,000 target?"
        python main.py            # interactive mode, type 'exit' to quit
    """
    import argparse
    
    from agent import build_agent
    from config import check_keys
    from output import print_trace, text_of
    
    
    def ask(agent, question: str, trace: bool) -> None:
        response = agent.invoke({"messages": [{"role": "user", "content": question}]})
        if trace:
            print_trace(response["messages"])
        print("\n" + "=" * 70)
        print(text_of(response["messages"][-1].content).strip())
        print("=" * 70)
    
    
    def main() -> None:
        parser = argparse.ArgumentParser(description="Gemini + Tavily research agent")
        parser.add_argument("question", nargs="*", help="Your question (leave empty for interactive mode)")
        parser.add_argument("--trace", action="store_true", help="Show each tool call the agent makes")
        args = parser.parse_args()
    
        check_keys()
        agent = build_agent()
    
        if args.question:
            ask(agent, " ".join(args.question), args.trace)
            return
    
        print("Research agent ready. Type a question, or 'exit' to quit.")
        while True:
            question = input("\nYou: ").strip()
            if question.lower() in {"exit", "quit", ""}:
                break
            ask(agent, question, args.trace)
    
    
    if __name__ == "__main__":
        main()
    Try these
    python main.py --trace "What is the latest iPhone model and its price in USD?"
    
    python main.py --trace "A 50-day joint replacement episode costs: hospital $21,500, post-acute care $8,400, readmission $3,200. The target price is $31,000. Are we above or below target, and by how much?"
    
    python main.py          # interactive mode

    What --trace prints (the format; your queries and results will differ):

    Example trace
    [You] What is the latest iPhone model and its price in USD?
    [Gemini -> tool] tavily_search({'query': 'latest iPhone model price USD'})
    [Tool result]    {"query": "latest iPhone model price USD", "results": [{"url": "https://www.newsnationnow.com/...
    [Gemini] final answer written
  8. Run the tests

    tests/test_tools.py
    """Run with: pytest -q   (no API keys or internet needed)"""
    import sys
    from pathlib import Path
    
    import pytest
    
    sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
    
    from output import text_of  # noqa: E402
    from tools import calculator, safe_calculate  # noqa: E402
    
    
    def test_bundle_cost_from_workshop():
        assert safe_calculate("21500 + 8400 + 3200") == 33100
        assert safe_calculate("21500 + 8400 + 3200 - 31000") == 2100
    
    
    def test_commas_and_brackets():
        assert safe_calculate("(1,199 - 999) * 2") == 400
    
    
    def test_rejects_code():
        with pytest.raises(ValueError):
            safe_calculate("__import__('os').system('dir')")
    
    
    def test_calculator_tool_output():
        assert calculator.invoke({"expression": "10 / 4"}) == "10 / 4 = 2.50"
        assert "Could not calculate" in calculator.invoke({"expression": "1/0"})
    
    
    def test_text_of_handles_gemini_blocks():
        assert text_of([{"type": "text", "text": "Hello"}]) == "Hello"
        assert text_of("Hello") == "Hello"
    Expected result
    pytest -q
    .....                                                                    [100%]
    5 passed

    These tests need no keys or internet, so they run in a second. They check the exact cost case from Test B, and that the calculator refuses code. Add a test every time you add a tool.

Add a new tool in 3 steps: write a function with a clear docstring and @tool, add it to build_tools(), and add a test.

Example: currency converter
# tools.py - example: add a currency converter in 3 steps
@tool
def convert_currency(amount: float, rate: float, from_code: str, to_code: str) -> str:
    """Convert an amount between currencies. Search for today's exchange rate first."""
    return f"{amount:,.2f} {from_code} = {amount * rate:,.2f} {to_code} (rate {rate})"

def build_tools(max_results: int) -> list:
    web_search = TavilySearch(max_results=max_results)
    return [web_search, calculator, today, convert_currency]   # step 2: register it
Developers

6. Problems and fixes

ProblemCauseFix
NameError: name 'llm' is not definedNotebook cells ran out of orderRun cells top to bottom, or use the project (main.py builds everything in order)
Missing GOOGLE_API_KEY / Tavily 401No .env, wrong folder or a typo in the key nameCopy .env.example to .env in the project folder
Output is [{'type': 'text', 'text': ...}]Gemini 3.x returns content blockstext_of() in output.py
LaTeX such as \mathbf{...} in answersThe model formats maths for a renderer you don't havePlain-text rule in the system prompt
Raw search results full of SVG or page codeSome pages (here, a social post) return markupNormal; the model ignores most of it. Use include_domains or exclude_domains in TavilySearch to prefer trusted sites
Answer contradicts the sourcesHallucinationGrounding rules, --trace, human check
Tavily stops working mid-monthFree credits used up (1 per basic search)Lower SEARCH_MAX_RESULTS, cache answers, check usage at app.tavily.com
404 model / 503 high demandRetired model name / busy serversChange GEMINI_MODEL in .env; retries are already on
EveryoneDevelopers

7. Ten POC ideas to build next

Each idea reuses this project. You keep agent.py and change the tools and the rules. They're ordered from easiest to hardest, and each one makes a good portfolio or interview project.

Daily news brief

Starter

"Summarise today's top 5 AI news stories with links." Run it every morning from a scheduled task.

+ Tavily topic="news", time_range="day"

Price comparison assistant

Starter

Compare a product across 3 stores and convert currencies. Good practice for grounding and showing sources.

+ convert_currency tool (step 8 example)

Company briefing for sales calls

Starter

Give it a company name and get funding, leadership changes, recent news and talking points on one page.

+ rules for a fixed output template

Healthcare episode cost checker

Intermediate

Extends Test B: read episode costs from a CSV, compare each one with its target price, and flag the episodes over target with reasons.

+ read_csv tool, calculator, no web search

Job market analyser

Intermediate

"Which skills appear most in senior data engineer jobs in Bengaluru?" Search, extract skills and count them.

+ structured output (response_format)

Ask your PDFs (RAG)

Intermediate

Answer questions from your own policy or course PDFs with page citations. The most requested enterprise pattern.

+ retriever tool: embeddings + Chroma or FAISS

Talk to a database

Intermediate

Ask plain-English questions about a sample sales database. The agent writes and runs read-only SQL.

+ SQLite read-only query tool

Support ticket triage

Advanced

Classify incoming tickets, search the help centre for an answer, and draft a reply for a person to approve.

+ classifier output, human-approval step

Chat with memory

Advanced

Follow-up questions such as "and in euros?" work because the agent remembers the conversation.

+ InMemorySaver checkpointer + thread_id

Web app version

Advanced

Put the agent behind a Streamlit or Gradio page with a chat box and a "show sources" panel, then share it.

+ app.py calling build_agent()

For every POC, show three things: it works (a demo), you can see why (the trace), and you know where it fails (one honest example, like Test A). That third point is what makes a POC convincing to a manager or an interviewer.

Everyone

8. Practice and quiz

Practice tasks

  1. Run Test A with --trace. Does the agent's price now match the search results? Note which URLs it cites.
  2. Change SEARCH_MAX_RESULTS to 2 and to 8. How do the answer quality and the number of credits used change?
  3. Add the convert_currency tool and ask for the iPhone 18 Pro price in GBP and INR.
  4. Pick one POC from section 7 and build the first version in under an hour.

Quick quiz

1. Our agent searched, found "$1,199", and answered "$1,439". What is this called?

Hallucination: the model produced a fact that its sources do not support, even though it had the right data.

2. Why add a calculator when Gemini got $33,100 right?

Models do maths by predicting text, so they are sometimes wrong. A tool is exact every time and leaves a record in the trace.

3. Why must the calculator not use eval()?

eval() runs any Python code. Text from the model, or from a web page it read, could then run commands on your computer.

4. What caused NameError: name 'llm' is not defined?

A notebook cell used llm before the cell that created it had run. The project version avoids this by building everything in order.

5. How many Tavily searches does the free plan cover?

1,000 credits a month. A basic search costs 1 credit and an advanced search costs 2.

Summary

Day 3 swaps DuckDuckGo for Tavily, adds a calculator and a date tool, and gives the agent written rules: search for current facts, use only what the sources say, do maths with the calculator, cite URLs and write plain text. The notebook becomes a small project with one job per file, pinned versions, a command line with --trace, and tests that run without keys.

The most important lesson came from a run without errors: the agent found the right price and still reported a wrong one. Build agents that show their sources and their steps, and check them.