Gemini thinks, Tavily searches, a calculator does the maths, and the agent lists its sources. We turn a Jupyter notebook into a small project you can run, test and extend on your own laptop, then look at real results to see where agents still go wrong.
In Day 2 our agent searched the web with DuckDuckGo. It worked, but it sometimes failed with network errors and returned short, messy snippets. Today we switch to Tavily, a search service built for AI agents. It returns cleaner text, a relevance score and the source URL for each result. It needs a free API key (1,000 searches a month, no card).
We also add two small tools, a calculator and today's date, plus written rules that tell the agent to cite its sources. Finally, we move the code out of one long notebook into a few short files, so it is easier to read, test and change.
| Piece | Older LangChain tutorials | This workshop |
|---|---|---|
| Model | OpenAI ChatOpenAI | Google Gemini ChatGoogleGenerativeAI (free tier) |
| Search | Serper / DuckDuckGo | Tavily TavilySearch |
| Loading tools | load_tools() | A plain Python list: [search, calculator, today] |
| Agent | ReAct PromptTemplate + AgentExecutor | create_agent() with native tool calling |
| Run | agent_executor.run() | agent.invoke({"messages": [...]}) |
| Code layout | One notebook | Notebook to learn, small project to run and maintain |
We ran the notebook version (Gemini + Tavily, no system prompt and no calculator) on two questions in October 2026. Both runs finished without errors. One answer was right, and one contradicted its own search results.
| Model | What the search results said | What the agent answered |
|---|---|---|
| iPhone 18 Pro | from $1,199 (NewsNation, CNET) | from $1,439 |
| iPhone 18 Pro Max | from $1,299 | "$1,299 to $1,399" |
| iPhone Duo (foldable) | from $1,999, on sale 23 Oct | around $1,999 |
The agent searched, found the right numbers, and still gave a wrong price. Searching does not guarantee correctness. The model can still blend the results with its own guesses, which is called hallucination. In a shopping or finance tool, a $240 error like this matters.
"List the top 3 Medicare providers. A 50-day joint replacement episode costs hospital $21,500, post-acute care $8,400 and readmission $3,200. The target price is $31,000. Are we above or below target?"
Total $33,100, so $2,100 above target. The maths is correct. It also noted that "Medicare providers" was ambiguous and answered with the largest Medicare Advantage insurers (UnitedHealthcare, Humana, CVS Health/Aetna).
The model did the maths in its head, which happened to be right this time. The output was full of LaTeX code such as \mathbf{\$33,100}, and it gave no source links for the market-share figures.
| Problem we saw | Fix | Where in the code |
|---|---|---|
| Price contradicted the sources | A rule: use only facts that appear in the search results, and say when sources disagree. List the source URLs. | agent.py SYSTEM_PROMPT rules 2 and 4 |
| Mental maths | A calculator tool, plus a rule to use it for every calculation | tools.py calculator |
| LaTeX in the output | A rule: plain text and simple bullets only | agent.py rule 5 |
| Hard to see what happened | --trace prints every tool call and result | output.py print_trace |
Rules reduce these mistakes but do not remove them. For anything important, compare the answer with the sources in the trace, or keep a person in the loop.
Notebooks are great for learning, but they get messy. Our Day 3 notebook hit NameError: name 'llm' is not defined because a cell ran before the cell that created llm. A project with small files fixes that: each file has one job, settings live in one place, and you can run it with one command and test it without API keys.
tavily_agent/
├── .env.example ← copy to .env, add keys
├── .gitignore ← keeps .env out of Git
├── requirements.txt ← pinned versions
├── config.py ← model, limits, key check
├── tools.py ← search, calculator, today
├── agent.py ← Gemini + tools + rules
├── output.py ← clean text, trace
├── main.py ← command line
├── README.md
└── tests/
└── test_tools.py| Use | When |
|---|---|
| Notebook | Learning, trying ideas, demos in a workshop |
| Project | Running again next month, sharing on GitHub, adding tools, interviews and POCs |
# Windows PowerShell (macOS/Linux: python3, source .venv/bin/activate, cp)
cd tavily_agent
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
copy .env.example .env # open .env and paste your two keys# Versions tested in the workshop notebook (Oct 2026). Pin them so the project keeps working.
langchain==1.4.3
langchain-core==1.6.7
langgraph==1.2.14
langchain-google-genai==4.4.0
langchain-tavily==0.2.18
python-dotenv==1.2.4
pytest>=8Pinned versions are the ones the workshop notebook ran with. Pinning means the project still works in three months, even after the libraries change.
# Copy this file to ".env" and paste your real keys. Never commit .env to GitHub.
GOOGLE_API_KEY=your_google_api_key # https://aistudio.google.com/apikey
TAVILY_API_KEY=your_tavily_api_key # https://app.tavily.com (free: 1,000 credits/month)
# Optional settings
GEMINI_MODEL=gemini-3.5-flash-lite
SEARCH_MAX_RESULTS=5Get the Gemini key from Google AI Studio and the Tavily key from app.tavily.com. .gitignore already stops .env from being committed.
"""All settings in one place. Change the model here or in .env, never inside the code."""
import os
from dataclasses import dataclass
from dotenv import load_dotenv
load_dotenv()
REQUIRED_KEYS = ("GOOGLE_API_KEY", "TAVILY_API_KEY")
@dataclass(frozen=True)
class Settings:
model: str = os.getenv("GEMINI_MODEL", "gemini-3.5-flash-lite")
search_max_results: int = int(os.getenv("SEARCH_MAX_RESULTS", "5"))
timeout_seconds: int = 60
max_retries: int = 3
def check_keys() -> None:
"""Stop early with a clear message instead of a long stack trace later."""
missing = [key for key in REQUIRED_KEYS if not os.getenv(key)]
if missing:
raise SystemExit(
f"Missing {', '.join(missing)}. Copy .env.example to .env and add your keys."
)When a model name stops working (Day 2 hit two model-name errors), you change one line in .env, not code in five places. check_keys() stops early with a clear message.
from dotenv import load_dotenv
from langchain_tavily import TavilySearch
load_dotenv()
search = TavilySearch(max_results=5)
result = search.invoke({"query": "latest iPhone model and price in USD"})
for item in result["results"]:
print(f"{item['score']:.2f} {item['title']}\n {item['url']}")Results from our run:
0.72 Apple - Explore the latest iPhone lineup, with prices...
https://www.facebook.com/apple/posts/...
0.64 Apple raises iPhone prices: Here's what each model costs now
https://www.newsnationnow.com/business/your-money/apple-raises-iphone-prices
0.48 Best iPhone in 2026: Apple's Phones Cost More, Here Are the Ones to Buy - CNET
https://www.cnet.com/tech/mobile/best-iphone
0.44 Apple's iPhone Pricing is a MESS
https://www.youtube.com/watch?v=hSfpZKA6Pzg
0.44 Which iPhone Should You Buy? - Consumer Reports
https://www.consumerreports.org/electronics-computers/cell-phones/which-iphone-should-you-buy-a7028071026Testing a tool alone before adding it to an agent tells you whether a problem is in the search or in the model. Notice that the top-scored result was a social media post, so a high score doesn't always mean the best source.
"""Tools the agent may call. Each docstring is what Gemini reads to decide when to use it."""
import ast
import operator
from datetime import date
from langchain.tools import tool
from langchain_tavily import TavilySearch
_OPERATORS = {
ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,
ast.Div: operator.truediv, ast.Pow: operator.pow, ast.Mod: operator.mod,
ast.USub: operator.neg, ast.UAdd: operator.pos,
}
def _evaluate(node: ast.AST) -> float:
if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):
return node.value
if isinstance(node, ast.BinOp) and type(node.op) in _OPERATORS:
return _OPERATORS[type(node.op)](_evaluate(node.left), _evaluate(node.right))
if isinstance(node, ast.UnaryOp) and type(node.op) in _OPERATORS:
return _OPERATORS[type(node.op)](_evaluate(node.operand))
raise ValueError("Only numbers and + - * / % ** ( ) are allowed")
def safe_calculate(expression: str) -> float:
"""Evaluate arithmetic without eval(), so the model cannot run arbitrary code."""
return _evaluate(ast.parse(expression.replace(",", ""), mode="eval").body)
@tool
def calculator(expression: str) -> str:
"""Do exact arithmetic, e.g. '21500 + 8400 + 3200 - 31000'. Use this for every calculation."""
try:
return f"{expression} = {safe_calculate(expression):,.2f}"
except (ValueError, SyntaxError, ZeroDivisionError) as err:
return f"Could not calculate '{expression}': {err}"
@tool
def today() -> str:
"""Return today's date. Use it when the question depends on 'latest', 'current' or 'this year'."""
return date.today().isoformat()
def build_tools(max_results: int) -> list:
web_search = TavilySearch(max_results=max_results) # reads TAVILY_API_KEY from the environment
return [web_search, calculator, today]The calculator parses the expression with Python's ast module and allows only numbers and operators. Never pass model output to eval(): a model, or a web page it read, could make it run any code on your laptop.
"""Builds the agent: Gemini (the brain) + tools (the hands) + rules (the system prompt)."""
from langchain.agents import create_agent
from langchain_google_genai import ChatGoogleGenerativeAI
from config import Settings
from tools import build_tools
SYSTEM_PROMPT = """You are a careful research assistant.
Rules:
1. For anything that changes over time (products, prices, rankings, news, laws), use web search. Do not answer from memory.
2. Use only facts that appear in the search results. If sources disagree or a fact is missing, say so.
3. Use the calculator tool for every calculation, even simple ones.
4. End with a 'Sources:' list of the URLs you used.
5. Write plain text with simple bullet points. Do not use LaTeX or math formatting."""
def build_agent(settings: Settings | None = None):
settings = settings or Settings()
llm = ChatGoogleGenerativeAI(
model=settings.model,
timeout=settings.timeout_seconds,
max_retries=settings.max_retries,
)
return create_agent(
model=llm,
tools=build_tools(settings.search_max_results),
system_prompt=SYSTEM_PROMPT,
)The system prompt turns what we learned in section 3 into instructions. Each rule fixes one problem we actually saw.
"""Helpers to turn the agent's raw response into readable text."""
def text_of(content) -> str:
"""Gemini 3.x returns a list of content blocks; older models return a string."""
if isinstance(content, list):
return "".join(part.get("text", "") for part in content if isinstance(part, dict))
return str(content)
def print_trace(messages) -> None:
"""Show every step of the agent loop: question, tool calls, tool results, answer."""
for msg in messages:
kind = type(msg).__name__
if kind == "HumanMessage":
print(f"\n[You] {text_of(msg.content).strip()[:200]}")
elif kind == "AIMessage" and msg.tool_calls:
for call in msg.tool_calls:
print(f"[Gemini -> tool] {call['name']}({call['args']})")
elif kind == "ToolMessage":
print(f"[Tool result] {text_of(msg.content)[:160].strip()}...")
elif kind == "AIMessage":
print("[Gemini] final answer written")"""Run the agent from the terminal.
python main.py "What is the latest iPhone model and its price in USD?"
python main.py --trace "Is a $33,100 episode above a $31,000 target?"
python main.py # interactive mode, type 'exit' to quit
"""
import argparse
from agent import build_agent
from config import check_keys
from output import print_trace, text_of
def ask(agent, question: str, trace: bool) -> None:
response = agent.invoke({"messages": [{"role": "user", "content": question}]})
if trace:
print_trace(response["messages"])
print("\n" + "=" * 70)
print(text_of(response["messages"][-1].content).strip())
print("=" * 70)
def main() -> None:
parser = argparse.ArgumentParser(description="Gemini + Tavily research agent")
parser.add_argument("question", nargs="*", help="Your question (leave empty for interactive mode)")
parser.add_argument("--trace", action="store_true", help="Show each tool call the agent makes")
args = parser.parse_args()
check_keys()
agent = build_agent()
if args.question:
ask(agent, " ".join(args.question), args.trace)
return
print("Research agent ready. Type a question, or 'exit' to quit.")
while True:
question = input("\nYou: ").strip()
if question.lower() in {"exit", "quit", ""}:
break
ask(agent, question, args.trace)
if __name__ == "__main__":
main()python main.py --trace "What is the latest iPhone model and its price in USD?"
python main.py --trace "A 50-day joint replacement episode costs: hospital $21,500, post-acute care $8,400, readmission $3,200. The target price is $31,000. Are we above or below target, and by how much?"
python main.py # interactive modeWhat --trace prints (the format; your queries and results will differ):
[You] What is the latest iPhone model and its price in USD?
[Gemini -> tool] tavily_search({'query': 'latest iPhone model price USD'})
[Tool result] {"query": "latest iPhone model price USD", "results": [{"url": "https://www.newsnationnow.com/...
[Gemini] final answer written"""Run with: pytest -q (no API keys or internet needed)"""
import sys
from pathlib import Path
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from output import text_of # noqa: E402
from tools import calculator, safe_calculate # noqa: E402
def test_bundle_cost_from_workshop():
assert safe_calculate("21500 + 8400 + 3200") == 33100
assert safe_calculate("21500 + 8400 + 3200 - 31000") == 2100
def test_commas_and_brackets():
assert safe_calculate("(1,199 - 999) * 2") == 400
def test_rejects_code():
with pytest.raises(ValueError):
safe_calculate("__import__('os').system('dir')")
def test_calculator_tool_output():
assert calculator.invoke({"expression": "10 / 4"}) == "10 / 4 = 2.50"
assert "Could not calculate" in calculator.invoke({"expression": "1/0"})
def test_text_of_handles_gemini_blocks():
assert text_of([{"type": "text", "text": "Hello"}]) == "Hello"
assert text_of("Hello") == "Hello"pytest -q
..... [100%]
5 passedThese tests need no keys or internet, so they run in a second. They check the exact cost case from Test B, and that the calculator refuses code. Add a test every time you add a tool.
Add a new tool in 3 steps: write a function with a clear docstring and @tool, add it to build_tools(), and add a test.
# tools.py - example: add a currency converter in 3 steps
@tool
def convert_currency(amount: float, rate: float, from_code: str, to_code: str) -> str:
"""Convert an amount between currencies. Search for today's exchange rate first."""
return f"{amount:,.2f} {from_code} = {amount * rate:,.2f} {to_code} (rate {rate})"
def build_tools(max_results: int) -> list:
web_search = TavilySearch(max_results=max_results)
return [web_search, calculator, today, convert_currency] # step 2: register it| Problem | Cause | Fix |
|---|---|---|
NameError: name 'llm' is not defined | Notebook cells ran out of order | Run cells top to bottom, or use the project (main.py builds everything in order) |
Missing GOOGLE_API_KEY / Tavily 401 | No .env, wrong folder or a typo in the key name | Copy .env.example to .env in the project folder |
Output is [{'type': 'text', 'text': ...}] | Gemini 3.x returns content blocks | text_of() in output.py |
LaTeX such as \mathbf{...} in answers | The model formats maths for a renderer you don't have | Plain-text rule in the system prompt |
| Raw search results full of SVG or page code | Some pages (here, a social post) return markup | Normal; the model ignores most of it. Use include_domains or exclude_domains in TavilySearch to prefer trusted sites |
| Answer contradicts the sources | Hallucination | Grounding rules, --trace, human check |
| Tavily stops working mid-month | Free credits used up (1 per basic search) | Lower SEARCH_MAX_RESULTS, cache answers, check usage at app.tavily.com |
| 404 model / 503 high demand | Retired model name / busy servers | Change GEMINI_MODEL in .env; retries are already on |
Each idea reuses this project. You keep agent.py and change the tools and the rules. They're ordered from easiest to hardest, and each one makes a good portfolio or interview project.
"Summarise today's top 5 AI news stories with links." Run it every morning from a scheduled task.
+ Tavily topic="news", time_range="day"Compare a product across 3 stores and convert currencies. Good practice for grounding and showing sources.
+ convert_currency tool (step 8 example)Give it a company name and get funding, leadership changes, recent news and talking points on one page.
+ rules for a fixed output templateExtends Test B: read episode costs from a CSV, compare each one with its target price, and flag the episodes over target with reasons.
+ read_csv tool, calculator, no web search"Which skills appear most in senior data engineer jobs in Bengaluru?" Search, extract skills and count them.
+ structured output (response_format)Answer questions from your own policy or course PDFs with page citations. The most requested enterprise pattern.
+ retriever tool: embeddings + Chroma or FAISSAsk plain-English questions about a sample sales database. The agent writes and runs read-only SQL.
+ SQLite read-only query toolClassify incoming tickets, search the help centre for an answer, and draft a reply for a person to approve.
+ classifier output, human-approval stepFollow-up questions such as "and in euros?" work because the agent remembers the conversation.
+ InMemorySaver checkpointer + thread_idPut the agent behind a Streamlit or Gradio page with a chat box and a "show sources" panel, then share it.
+ app.py calling build_agent()For every POC, show three things: it works (a demo), you can see why (the trace), and you know where it fails (one honest example, like Test A). That third point is what makes a POC convincing to a manager or an interviewer.
--trace. Does the agent's price now match the search results? Note which URLs it cites.SEARCH_MAX_RESULTS to 2 and to 8. How do the answer quality and the number of credits used change?convert_currency tool and ask for the iPhone 18 Pro price in GBP and INR.Hallucination: the model produced a fact that its sources do not support, even though it had the right data.
Models do maths by predicting text, so they are sometimes wrong. A tool is exact every time and leaves a record in the trace.
eval() runs any Python code. Text from the model, or from a web page it read, could then run commands on your computer.
A notebook cell used llm before the cell that created it had run. The project version avoids this by building everything in order.
1,000 credits a month. A basic search costs 1 credit and an advanced search costs 2.
Day 3 swaps DuckDuckGo for Tavily, adds a calculator and a date tool, and gives the agent written rules: search for current facts, use only what the sources say, do maths with the calculator, cite URLs and write plain text. The notebook becomes a small project with one job per file, pinned versions, a command line with --trace, and tests that run without keys.
The most important lesson came from a run without errors: the agent found the right price and still reported a wrong one. Build agents that show their sources and their steps, and check them.