{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
""
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Google GenAI"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"In this notebook, we show how to use the `google-genai` Python SDK with LlamaIndex to interact with Google GenAI models.\n",
"\n",
"If you're opening this Notebook on colab, you will need to install LlamaIndex 🦙 and the `google-genai` Python SDK."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"%pip install llama-index-llms-google-genai llama-index"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Basic Usage\n",
"\n",
"You will need to get an API key from [Google AI Studio](https://makersuite.google.com/app/apikey). Once you have one, you can either pass it explicity to the model, or use the `GOOGLE_API_KEY` environment variable."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"import os\n",
"\n",
"os.environ[\"GOOGLE_API_KEY\"] = \"...\""
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Basic Usage\n",
"\n",
"You can call `complete` with a prompt:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Paul Graham is a prominent figure in the tech world, best known for his work as a programmer, essayist, and venture capitalist. Here's a breakdown of his key contributions:\n",
"\n",
"* **Programmer and Hacker:** He's a skilled programmer, particularly in Lisp. He co-founded Viaweb, which was one of the first software-as-a-service (SaaS) companies, providing tools for building online stores. Yahoo acquired Viaweb in 1998, and it became Yahoo! Store.\n",
"\n",
"* **Essayist:** Graham is a prolific and influential essayist. His essays cover a wide range of topics, including startups, programming, design, and societal trends. His writing style is known for being clear, concise, and thought-provoking. Many of his essays are considered essential reading for entrepreneurs and those interested in technology.\n",
"\n",
"* **Venture Capitalist and Founder of Y Combinator:** Perhaps his most significant contribution is co-founding Y Combinator (YC) in 2005. YC is a highly successful startup accelerator that provides seed funding, mentorship, and networking opportunities to early-stage startups. YC has funded many well-known companies, including Airbnb, Dropbox, Reddit, Stripe, and many others. Graham stepped down from his day-to-day role at YC in 2014 but remains involved.\n",
"\n",
"In summary, Paul Graham is a multifaceted individual who has made significant contributions to the tech industry as a programmer, essayist, and venture capitalist. He is particularly known for his role in founding and shaping Y Combinator, one of the world's leading startup accelerators.\n"
]
}
],
"source": [
"from llama_index.llms.google_genai import GoogleGenAI\n",
"\n",
"llm = GoogleGenAI(\n",
" model=\"gemini-2.5-flash\",\n",
" # api_key=\"some key\", # uses GOOGLE_API_KEY env var by default\n",
")\n",
"\n",
"resp = llm.complete(\"Who is Paul Graham?\")\n",
"print(resp)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"You can also call `chat` with a list of chat messages:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"assistant: Ahoy there, matey! Gather 'round, ye landlubbers, and listen to a tale that'll shiver yer timbers and curl yer toes! This be the story of One-Eyed Jack's Lost Parrot and the Great Mango Mayhem!\n",
"\n",
"Now, One-Eyed Jack, bless his barnacle-encrusted heart, was a fearsome pirate, alright. He could bellow louder than a hurricane, swing a cutlass like a dervish, and drink rum like a fish. But he had a soft spot, see? A soft spot for his parrot, Polly. Polly wasn't just any parrot, mind ye. She could mimic the captain's every cuss word, predict the weather by the way she ruffled her feathers, and had a particular fondness for shiny trinkets.\n",
"\n",
"One day, we were anchored off the coast of Mango Island, a lush paradise overflowing with the juiciest, sweetest mangoes ye ever did see. Jack, bless his greedy soul, decided we needed a cargo hold full of 'em. \"For scurvy prevention!\" he declared, winking with his good eye. More like for his own personal mango-eating contest, if ye ask me.\n",
"\n",
"We stormed ashore, cutlasses gleaming, ready to plunder the mango groves. But Polly, the little feathered devil, decided she'd had enough of the ship. She squawked, \"Shiny! Shiny!\" and took off like a green streak towards the heart of the island.\n",
"\n",
"Jack went ballistic! \"Polly! Polly, ye feathered fiend! Get back here!\" He chased after her, bellowing like a lovesick walrus. The rest of us, well, we were left to pick mangoes and try not to laugh ourselves silly.\n",
"\n",
"Now, Mango Island wasn't just full of mangoes. It was also home to a tribe of mischievous monkeys, the Mango Marauders, they were called. They were notorious for their love of pranks and their uncanny ability to steal anything that wasn't nailed down.\n",
"\n",
"Turns out, Polly had landed right in the middle of their territory. And those monkeys, they took one look at her shiny feathers and decided she was the perfect addition to their collection of stolen treasures. They snatched her up, chattering and screeching, and whisked her away to their hidden lair, a giant mango tree hollowed out by time.\n",
"\n",
"Jack, bless his stubborn heart, followed the sound of Polly's squawks. He hacked through vines, dodged falling mangoes, and even wrestled a particularly grumpy iguana, all in pursuit of his feathered friend.\n",
"\n",
"Finally, he reached the mango tree. He peered inside and saw Polly, surrounded by a horde of monkeys, all admiring her shiny feathers. And Polly? She was having the time of her life, mimicking the monkeys' chattering and stealing their mangoes!\n",
"\n",
"Jack, instead of getting angry, started to laugh. A hearty, booming laugh that shook the very foundations of the tree. The monkeys, startled, dropped their mangoes and stared at him.\n",
"\n",
"Then, Polly, seeing her captain, squawked, \"Rum! Rum for everyone!\"\n",
"\n",
"And that, me hearties, is how One-Eyed Jack ended up sharing a barrel of rum with a tribe of mango-loving monkeys. We spent the rest of the day feasting on mangoes, drinking rum, and listening to Polly mimic the monkeys' antics. We even managed to fill the cargo hold with mangoes, though I suspect a good portion of them were already half-eaten by the monkeys.\n",
"\n",
"So, the moral of the story, me lads? Even the fiercest pirate has a soft spot, and sometimes, the best treasures are the ones you least expect. And always, ALWAYS, keep an eye on yer parrot! Now, who's for another round of grog?\n",
"\n"
]
}
],
"source": [
"from llama_index.core.llms import ChatMessage\n",
"from llama_index.llms.google_genai import GoogleGenAI\n",
"\n",
"messages = [\n",
" ChatMessage(\n",
" role=\"system\", content=\"You are a pirate with a colorful personality\"\n",
" ),\n",
" ChatMessage(role=\"user\", content=\"Tell me a story\"),\n",
"]\n",
"llm = GoogleGenAI(model=\"gemini-2.5-flash\")\n",
"resp = llm.chat(messages)\n",
"\n",
"print(resp)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Streaming Support\n",
"\n",
"Every method supports streaming through the `stream_` prefix."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Paul Graham is a prominent figure in the tech world, best known for his work as a computer programmer, essayist, venture capitalist, and co-founder of the startup accelerator Y Combinator. Here's a breakdown of his key accomplishments and contributions:\n",
"\n",
"* **Computer Programmer and Author:** Graham holds a Ph.D. in computer science from Harvard University. He is known for his work on Lisp, a programming language, and for developing Viaweb, one of the first software-as-a-service (SaaS) companies, which was later acquired by Yahoo! and became Yahoo! Store. He's also the author of several influential books on programming and entrepreneurship, including \"On Lisp,\" \"ANSI Common Lisp,\" \"Hackers & Painters,\" and \"A Plan for Spam.\"\n",
"\n",
"* **Essayist:** Graham is a prolific essayist, writing on a wide range of topics including technology, startups, art, philosophy, and society. His essays are known for their insightful observations, clear writing style, and often contrarian viewpoints. They are widely read and discussed in the tech community. You can find his essays on his website, paulgraham.com.\n",
"\n",
"* **Venture Capitalist and Y Combinator:** Graham co-founded Y Combinator (YC) in 2005 with Jessica Livingston, Robert Morris, and Trevor Blackwell. YC is a highly successful startup accelerator that provides seed funding, mentorship, and networking opportunities to early-stage startups. YC has funded many well-known companies, including Airbnb, Dropbox, Reddit, Stripe, and many others. While he stepped down from day-to-day operations at YC in 2014, his influence on the organization and the startup ecosystem remains significant.\n",
"\n",
"In summary, Paul Graham is a multifaceted individual who has made significant contributions to computer science, entrepreneurship, and the broader tech culture. He is highly regarded for his technical expertise, insightful writing, and his role in shaping the modern startup landscape."
]
}
],
"source": [
"from llama_index.llms.google_genai import GoogleGenAI\n",
"\n",
"llm = GoogleGenAI(model=\"gemini-2.5-flash\")\n",
"\n",
"resp = llm.stream_complete(\"Who is Paul Graham?\")\n",
"for r in resp:\n",
" print(r.delta, end=\"\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Paul Graham is a prominent figure in the tech world, best known for his work as a programmer, essayist, and venture capitalist. Here's a breakdown of his key contributions:\n",
"\n",
"* **Programmer and Hacker:** He is a skilled programmer, particularly in Lisp. He co-founded Viaweb, one of the first software-as-a-service (SaaS) companies, which was later acquired by Yahoo! and became Yahoo! Store.\n",
"\n",
"* **Essayist:** Graham is a prolific and influential essayist, writing on topics ranging from programming and startups to art, philosophy, and social commentary. His essays are known for their clarity, insight, and often contrarian viewpoints. They are widely read and discussed in the tech community.\n",
"\n",
"* **Venture Capitalist:** He co-founded Y Combinator (YC) in 2005, a highly successful startup accelerator. YC has funded and mentored numerous well-known companies, including Airbnb, Dropbox, Reddit, Stripe, and many others. Graham's approach to early-stage investing and startup mentorship has had a significant impact on the startup ecosystem.\n",
"\n",
"In summary, Paul Graham is a multifaceted individual who has made significant contributions to the tech industry as a programmer, essayist, and venture capitalist. He is particularly influential in the startup world through his work with Y Combinator."
]
}
],
"source": [
"from llama_index.core.llms import ChatMessage\n",
"\n",
"messages = [\n",
" ChatMessage(role=\"user\", content=\"Who is Paul Graham?\"),\n",
"]\n",
"\n",
"resp = llm.stream_chat(messages)\n",
"for r in resp:\n",
" print(r.delta, end=\"\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Async Usage\n",
"\n",
"Every synchronous method has an async counterpart."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Paul Graham is a prominent figure in the tech world, best known for his work as a programmer, essayist, and venture capitalist. Here's a breakdown of his key accomplishments and roles:\n",
"\n",
"* **Programmer and Hacker:** He holds a Ph.D. in computer science from Harvard and is known for his work on Lisp, a programming language. He co-founded Viaweb, one of the first software-as-a-service (SaaS) companies, which was later acquired by Yahoo! and became Yahoo! Store.\n",
"\n",
"* **Essayist:** Graham is a prolific and influential essayist, writing on topics ranging from programming and startups to art, philosophy, and social commentary. His essays are widely read and discussed in the tech community.\n",
"\n",
"* **Venture Capitalist:** He co-founded Y Combinator (YC) in 2005, a highly successful startup accelerator that has funded companies like Airbnb, Dropbox, Reddit, Stripe, and many others. YC provides seed funding, mentorship, and networking opportunities to early-stage startups. While he stepped back from day-to-day operations at YC in 2014, he remains a significant figure in the venture capital world.\n",
"\n",
"In summary, Paul Graham is a multifaceted individual who has made significant contributions to the fields of computer science, entrepreneurship, and venture capital. He is highly regarded for his insightful writing and his role in shaping the modern startup ecosystem."
]
}
],
"source": [
"from llama_index.llms.google_genai import GoogleGenAI\n",
"\n",
"llm = GoogleGenAI(model=\"gemini-2.5-flash\")\n",
"\n",
"resp = await llm.astream_complete(\"Who is Paul Graham?\")\n",
"async for r in resp:\n",
" print(r.delta, end=\"\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"assistant: Paul Graham is a prominent figure in the tech world, best known for his work as a programmer, essayist, and venture capitalist. Here's a breakdown of his key accomplishments and contributions:\n",
"\n",
"* **Programmer and Hacker:** He is a skilled programmer, particularly in Lisp. He co-founded Viaweb, one of the first software-as-a-service (SaaS) companies, which was later acquired by Yahoo! and became Yahoo! Store.\n",
"\n",
"* **Essayist:** Graham is a prolific and influential essayist, writing on topics ranging from programming and startups to art, design, and societal trends. His essays are known for their insightful observations, contrarian viewpoints, and clear writing style. Many of his essays are available on his website, paulgraham.com.\n",
"\n",
"* **Venture Capitalist and Y Combinator:** He co-founded Y Combinator (YC) in 2005, a highly successful startup accelerator that has funded numerous well-known companies, including Airbnb, Dropbox, Reddit, Stripe, and many others. YC provides seed funding, mentorship, and networking opportunities to early-stage startups. Graham played a key role in shaping YC's philosophy and approach to investing.\n",
"\n",
"* **Author:** He has written several books, including \"On Lisp\" and \"Hackers & Painters: Big Ideas from the Age of Enlightenment.\"\n",
"\n",
"In summary, Paul Graham is a multifaceted individual who has made significant contributions to the tech industry as a programmer, essayist, and venture capitalist. He is particularly influential in the startup world through his work with Y Combinator.\n"
]
}
],
"source": [
"messages = [\n",
" ChatMessage(role=\"user\", content=\"Who is Paul Graham?\"),\n",
"]\n",
"\n",
"resp = await llm.achat(messages)\n",
"print(resp)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Vertex AI Support\n",
"\n",
"By providing the `region` and `project_id` parameters (either through environment variables or directly), you can enable usage through Vertex AI."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# Set environment variables\n",
"!export GOOGLE_GENAI_USE_VERTEXAI=true\n",
"!export GOOGLE_CLOUD_PROJECT='your-project-id'\n",
"!export GOOGLE_CLOUD_LOCATION='us-central1'"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Paul Graham is a prominent figure in the tech and startup world, best known for his roles as:\n",
"\n",
"* **Co-founder of Y Combinator (YC):** This is arguably his most influential role. YC is a highly successful startup accelerator that has funded companies like Airbnb, Dropbox, Stripe, Reddit, and many others. Graham's approach to funding and mentoring startups has significantly shaped the startup ecosystem.\n",
"\n",
"* **Essayist and Programmer:** Before YC, Graham was a programmer and essayist. He's known for his insightful and often contrarian essays on a wide range of topics, including programming, startups, design, and societal trends. His essays are widely read and discussed in the tech community.\n",
"\n",
"* **Founder of Viaweb (later Yahoo! Store):** Graham founded Viaweb, one of the first application service providers, which allowed users to build and manage online stores. It was acquired by Yahoo! in 1998 and became Yahoo! Store.\n",
"\n",
"In summary, Paul Graham is a highly influential figure in the startup world, known for his role in creating Y Combinator, his insightful essays, and his earlier success as a programmer and entrepreneur.\n"
]
}
],
"source": [
"from llama_index.llms.google_genai import GoogleGenAI\n",
"\n",
"# or set the parameters directly\n",
"llm = GoogleGenAI(\n",
" model=\"gemini-2.5-flash\",\n",
" vertexai_config={\"project\": \"your-project-id\", \"location\": \"us-central1\"},\n",
" # you should set the context window to the max input tokens for the model\n",
" context_window=200000,\n",
" max_tokens=512,\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Cached Content Support\n",
"\n",
"Google GenAI supports cached content for improved performance and cost efficiency when reusing large contexts across multiple requests. This is particularly useful for RAG applications, document analysis, and multi-turn conversations with consistent context.\n",
"\n",
"#### Benefits\n",
"\n",
"- **Faster responses**\n",
"- **Cost savings** through reduced input token usage\n",
"- **Consistent context** across multiple queries\n",
"- **Perfect for document analysis** with large files\n",
"\n",
"#### Creating Cached Content\n",
"\n",
"First, create cached content using the Google GenAI SDK:\n"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from google import genai\n",
"from google.genai.types import CreateCachedContentConfig, Content, Part\n",
"import time\n",
"\n",
"client = genai.Client(api_key=\"your-api-key\")\n",
"\n",
"# For VertexAI\n",
"# client = genai.Client(\n",
"# http_options=HttpOptions(api_version=\"v1\"),\n",
"# project=\"your-project-id\",\n",
"# location=\"us-central1\",\n",
"# vertexai=\"True\"\n",
"# )"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Option 1: Upload Local Files"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"# Upload and process local PDF files\n",
"pdf_file = client.files.upload(file=\"./your_document.pdf\")\n",
"while pdf_file.state.name == \"PROCESSING\":\n",
" print(\"Waiting for PDF to be processed.\")\n",
" time.sleep(2)\n",
" pdf_file = client.files.get(name=pdf_file.name)\n",
"\n",
"# Create cache with uploaded file\n",
"cache = client.caches.create(\n",
" model=\"gemini-2.5-flash\",\n",
" config=CreateCachedContentConfig(\n",
" display_name=\"Document Analysis Cache\",\n",
" system_instruction=(\n",
" \"You are an expert document analyzer. Answer questions \"\n",
" \"based on the provided documents with accuracy and detail.\"\n",
" ),\n",
" contents=[pdf_file], # Direct file reference\n",
" ttl=\"3600s\", # Cache for 1 hour\n",
" ),\n",
")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Option 2: Multiple Files with Content Structure"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Cache created: projects/391.../locations/us-central1/cachedContents/267...\n",
"Cached tokens: 43102\n"
]
}
],
"source": [
"# For multiple files or Cloud Storage files with VertexAI\n",
"contents = [\n",
" Content(\n",
" role=\"user\",\n",
" parts=[\n",
" Part.from_uri(\n",
" # file_uri=pdf_file.uri, # you can use the uploaded file's URI too\n",
" file_uri=\"gs://cloud-samples-data/generative-ai/pdf/2312.11805v3.pdf\",\n",
" mime_type=\"application/pdf\",\n",
" ),\n",
" Part.from_uri(\n",
" file_uri=\"gs://cloud-samples-data/generative-ai/pdf/2403.05530.pdf\",\n",
" mime_type=\"application/pdf\",\n",
" ),\n",
" ],\n",
" )\n",
"]\n",
"\n",
"cache = client.caches.create(\n",
" model=\"gemini-2.5-flash\",\n",
" config=CreateCachedContentConfig(\n",
" display_name=\"Multi-Document Cache\",\n",
" system_instruction=(\n",
" \"You are an expert researcher. Analyze and compare \"\n",
" \"information across the provided documents.\"\n",
" ),\n",
" contents=contents,\n",
" ttl=\"3600s\",\n",
" ),\n",
")\n",
"\n",
"print(f\"Cache created: {cache.name}\")\n",
"print(f\"Cached tokens: {cache.usage_metadata.total_token_count}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Using Cached Content with LlamaIndex\n",
"\n",
"Once you have created the cache, use it with LlamaIndex:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"assistant: Chapter 4, \"The Abstraction: The Process,\" introduces the concept of a process as a running program, which is a fundamental abstraction provided by the operating system (OS). Here are the key findings:\n",
"\n",
"1. **Process Definition:** A process is essentially a running program, characterized by its machine state, including memory (address space), registers (including the program counter and stack pointer), and I/O information.\n",
"\n",
"2. **Process API:** The OS provides a process API that includes functions for creating processes (Create), destroying processes (Destroy), waiting for processes to complete (Wait), controlling processes (Miscellaneous Control), and obtaining status information (Status).\n",
"\n",
"3. **Process Creation:** Creating a process involves loading code and static data into memory, allocating memory for the stack and heap, initializing the stack, and then starting the program at its entry point (main()).\n",
"\n",
"4. **Process States:** A process can be in one of three states: Running (executing on a processor), Ready (ready to run but not currently running), or Blocked (waiting for an event, such as I/O completion).\n",
"\n",
"5. **Data Structures:** The OS maintains data structures, such as a process list, to track the state of each process. These structures contain information like the register context (saved register values) and the process state.\n",
"\n",
"In essence, Chapter 4 lays the groundwork for understanding how the OS manages and virtualizes the CPU by introducing the concept of a process and its associated attributes and states.\n",
"\n"
]
}
],
"source": [
"from llama_index.llms.google_genai import GoogleGenAI\n",
"from llama_index.core.llms import ChatMessage\n",
"\n",
"llm = GoogleGenAI(\n",
" model=\"gemini-2.5-flash\",\n",
" api_key=\"your-api-key\",\n",
" cached_content=cache.name,\n",
")\n",
"\n",
"# For VertexAI\n",
"# llm = GoogleGenAI(\n",
"# model=\"gemini-2.5-flash\",\n",
"# vertexai_config={\"project\": \"your-project-id\", \"location\": \"us-central1\"},\n",
"# cached_content=cache.name\n",
"# )\n",
"\n",
"# Use the cached content\n",
"message = ChatMessage(\n",
" role=\"user\", content=\"Summarize the key findings from Chapter 4.\"\n",
")\n",
"response = llm.chat([message])\n",
"print(response)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Using Cached Content in Generation Config\n",
"\n",
"For request-level caching control:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Here are the first five chapters of the document, as listed in the Table of Contents:\n",
"\n",
"1. A Dialogue on the Book\n",
"2. Introduction to Operating Systems\n",
"3. A Dialogue on Virtualization\n",
"4. The Abstraction: The Process\n",
"5. Interlude: Process API\n"
]
}
],
"source": [
"import google.genai.types as types\n",
"\n",
"# Specify cached content per request\n",
"config = types.GenerateContentConfig(\n",
" cached_content=cache.name, temperature=0.1, max_output_tokens=1024\n",
")\n",
"\n",
"llm = GoogleGenAI(model=\"gemini-2.5-flash\", generation_config=config)\n",
"\n",
"response = llm.complete(\"List the first five chapters of the document\")\n",
"print(response)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Cache Management"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Cache: Document Analysis Cache (cachedContents/8v3va2x...)\n",
"Tokens: 77421\n",
"Created: 2025-07-08 16:06:11.821190+00:00\n",
"Expires: 2025-07-08 17:06:10.813310+00:00\n",
"Cache deleted\n"
]
}
],
"source": [
"# List all caches\n",
"caches = client.caches.list()\n",
"for cache_item in caches:\n",
" print(f\"Cache: {cache_item.display_name} ({cache_item.name})\")\n",
" print(f\"Tokens: {cache_item.usage_metadata.total_token_count}\")\n",
"\n",
"# Get cache details\n",
"cache_info = client.caches.get(name=cache.name)\n",
"print(f\"Created: {cache_info.create_time}\")\n",
"print(f\"Expires: {cache_info.expire_time}\")\n",
"\n",
"# Delete cache when done\n",
"client.caches.delete(name=cache.name)\n",
"print(\"Cache deleted\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Multi-Modal Support\n",
"\n",
"Using `ChatMessage` objects, you can pass in images and text to the LLM."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"--2025-03-14 10:59:00-- https://cdn.pixabay.com/photo/2021/12/12/20/00/play-6865967_640.jpg\n",
"Resolving cdn.pixabay.com (cdn.pixabay.com)... 104.18.40.96, 172.64.147.160\n",
"Connecting to cdn.pixabay.com (cdn.pixabay.com)|104.18.40.96|:443... connected.\n",
"HTTP request sent, awaiting response... 200 OK\n",
"Length: 71557 (70K) [binary/octet-stream]\n",
"Saving to: ‘image.jpg’\n",
"\n",
"image.jpg 100%[===================>] 69.88K --.-KB/s in 0.003s \n",
"\n",
"2025-03-14 10:59:00 (24.8 MB/s) - ‘image.jpg’ saved [71557/71557]\n",
"\n"
]
}
],
"source": [
"!wget https://cdn.pixabay.com/photo/2021/12/12/20/00/play-6865967_640.jpg -O image.jpg"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"assistant: The image contains four wooden dice with black dots on a dark gray surface. Each die shows a different number of dots, indicating different values.\n"
]
}
],
"source": [
"from llama_index.core.llms import ChatMessage, TextBlock, ImageBlock\n",
"from llama_index.llms.google_genai import GoogleGenAI\n",
"\n",
"llm = GoogleGenAI(model=\"gemini-2.5-flash\")\n",
"\n",
"messages = [\n",
" ChatMessage(\n",
" role=\"user\",\n",
" blocks=[\n",
" ImageBlock(path=\"image.jpg\", image_mimetype=\"image/jpeg\"),\n",
" TextBlock(text=\"What is in this image?\"),\n",
" ],\n",
" )\n",
"]\n",
"\n",
"resp = llm.chat(messages)\n",
"print(resp)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"You can also pass in documents."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"assistant: This research paper assesses and mitigates multi-turn jailbreak vulnerabilities in recent large language models (LLMs) using the Crescendo attack, evaluating prompt hardening and LLM-as-guardrail strategies across various task categories.\n",
"\n"
]
}
],
"source": [
"from llama_index.core.llms import DocumentBlock\n",
"\n",
"messages = [\n",
" ChatMessage(\n",
" role=\"user\",\n",
" blocks=[\n",
" DocumentBlock(\n",
" path=\"/path/to/your/test.pdf\",\n",
" document_mimetype=\"application/pdf\",\n",
" ),\n",
" TextBlock(text=\"Describe the document in a sentence.\"),\n",
" ],\n",
" )\n",
"]\n",
"\n",
"resp = llm.chat(messages)\n",
"print(resp)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"Finally, you can also pass videos."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"assistant: A white SpaceX Crew Dragon capsule is shown approaching and docking with a module of the International Space Station, with the Earth's curvature visible in the background.\n"
]
}
],
"source": [
"from llama_index.core.llms import VideoBlock\n",
"\n",
"messages = [\n",
" ChatMessage(\n",
" role=\"user\",\n",
" blocks=[\n",
" VideoBlock(\n",
" path=\"/path/to/your/video.mp4\", video_mimetype=\"video/mp4\"\n",
" ),\n",
" TextBlock(text=\"Describe this video in a sentence.\"),\n",
" ],\n",
" )\n",
"]\n",
"\n",
"resp = llm.chat(messages)\n",
"print(resp)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Structured Prediction\n",
"\n",
"LlamaIndex provides an intuitive interface for converting any LLM into a structured LLM through `structured_predict` - simply define the target Pydantic class (can be nested), and given a prompt, we extract out the desired object."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.llms.google_genai import GoogleGenAI\n",
"from llama_index.core.prompts import PromptTemplate\n",
"from llama_index.core.bridge.pydantic import BaseModel\n",
"from typing import List\n",
"\n",
"\n",
"class MenuItem(BaseModel):\n",
" \"\"\"A menu item in a restaurant.\"\"\"\n",
"\n",
" course_name: str\n",
" is_vegetarian: bool\n",
"\n",
"\n",
"class Restaurant(BaseModel):\n",
" \"\"\"A restaurant with name, city, and cuisine.\"\"\"\n",
"\n",
" name: str\n",
" city: str\n",
" cuisine: str\n",
" menu_items: List[MenuItem]\n",
"\n",
"\n",
"llm = GoogleGenAI(model=\"gemini-2.5-flash\")\n",
"prompt_tmpl = PromptTemplate(\n",
" \"Generate a restaurant in a given city {city_name}\"\n",
")\n",
"\n",
"# Option 1: Use `as_structured_llm`\n",
"restaurant_obj = (\n",
" llm.as_structured_llm(Restaurant)\n",
" .complete(prompt_tmpl.format(city_name=\"Miami\"))\n",
" .raw\n",
")\n",
"# Option 2: Use `structured_predict`\n",
"# restaurant_obj = llm.structured_predict(Restaurant, prompt_tmpl, city_name=\"Miami\")"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"name='Pasta Mia' city='Miami' cuisine='Italian' menu_items=[MenuItem(course_name='pasta', is_vegetarian=False)]\n"
]
}
],
"source": [
"print(restaurant_obj)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"#### Structured Prediction with Streaming\n",
"\n",
"Any LLM wrapped with `as_structured_llm` supports streaming through `stream_chat`."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"{'city': 'San Francisco',\n",
" 'cuisine': 'Italian',\n",
" 'menu_items': [{'course_name': 'pasta', 'is_vegetarian': False}],\n",
" 'name': 'Italian Delight'}\n"
]
},
{
"name": "stderr",
"output_type": "stream",
"text": [
"/var/folders/lw/xwsz_3yj4ln1gvkxhyddbvvw0000gn/T/ipykernel_76091/1885953561.py:11: PydanticDeprecatedSince20: The `dict` method is deprecated; use `model_dump` instead. Deprecated in Pydantic V2.0 to be removed in V3.0. See Pydantic V2 Migration Guide at https://errors.pydantic.dev/2.10/migration/\n",
" pprint(partial_output.raw.dict())\n"
]
},
{
"data": {
"text/plain": [
"Restaurant(name='Italian Delight', city='San Francisco', cuisine='Italian', menu_items=[MenuItem(course_name='pasta', is_vegetarian=False)])"
]
},
"execution_count": null,
"metadata": {},
"output_type": "execute_result"
}
],
"source": [
"from llama_index.core.llms import ChatMessage\n",
"from IPython.display import clear_output\n",
"from pprint import pprint\n",
"\n",
"input_msg = ChatMessage.from_str(\"Generate a restaurant in San Francisco\")\n",
"\n",
"sllm = llm.as_structured_llm(Restaurant)\n",
"stream_output = sllm.stream_chat([input_msg])\n",
"for partial_output in stream_output:\n",
" clear_output(wait=True)\n",
" pprint(partial_output.raw.dict())\n",
" restaurant_obj = partial_output.raw\n",
"\n",
"restaurant_obj"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Tool/Function Calling\n",
"\n",
"Google GenAI supports direct tool/function calling through the API. Using LlamaIndex, we can implement some core agentic tool calling patterns."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [],
"source": [
"from llama_index.core.tools import FunctionTool\n",
"from llama_index.core.llms import ChatMessage\n",
"from llama_index.llms.google_genai import GoogleGenAI\n",
"from datetime import datetime\n",
"\n",
"llm = GoogleGenAI(model=\"gemini-2.5-flash\")\n",
"\n",
"\n",
"def get_current_time(timezone: str) -> dict:\n",
" \"\"\"Get the current time\"\"\"\n",
" return {\n",
" \"time\": datetime.now().strftime(\"%Y-%m-%d %H:%M:%S\"),\n",
" \"timezone\": timezone,\n",
" }\n",
"\n",
"\n",
"# uses the tool name, any type annotations, and docstring to describe the tool\n",
"tool = FunctionTool.from_defaults(fn=get_current_time)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We can simply do a single pass to call the tool and get the result:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"{'time': '2025-03-14 10:59:05', 'timezone': 'America/New_York'}\n"
]
}
],
"source": [
"resp = llm.predict_and_call([tool], \"What is the current time in New York?\")\n",
"print(resp)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We can also use lower-level APIs to implement an agentic tool-calling loop!"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Calling get_current_time with {'timezone': 'America/New_York'}\n",
"Tool output: {'time': '2025-03-14 10:59:06', 'timezone': 'America/New_York'}\n",
"Final response: The current time in New York is 2025-03-14 10:59:06.\n"
]
}
],
"source": [
"chat_history = [\n",
" ChatMessage(role=\"user\", content=\"What is the current time in New York?\")\n",
"]\n",
"tools_by_name = {t.metadata.name: t for t in [tool]}\n",
"\n",
"resp = llm.chat_with_tools([tool], chat_history=chat_history)\n",
"tool_calls = llm.get_tool_calls_from_response(\n",
" resp, error_on_no_tool_call=False\n",
")\n",
"\n",
"if not tool_calls:\n",
" print(resp)\n",
"else:\n",
" while tool_calls:\n",
" # add the LLM's response to the chat history\n",
" chat_history.append(resp.message)\n",
"\n",
" for tool_call in tool_calls:\n",
" tool_name = tool_call.tool_name\n",
" tool_kwargs = tool_call.tool_kwargs\n",
"\n",
" print(f\"Calling {tool_name} with {tool_kwargs}\")\n",
" tool_output = tool.call(**tool_kwargs)\n",
" print(\"Tool output: \", tool_output)\n",
" chat_history.append(\n",
" ChatMessage(\n",
" role=\"tool\",\n",
" content=str(tool_output),\n",
" # most LLMs like Gemini, Anthropic, OpenAI, etc. need to know the tool call id\n",
" additional_kwargs={\"tool_call_id\": tool_call.tool_id},\n",
" )\n",
" )\n",
"\n",
" resp = llm.chat_with_tools([tool], chat_history=chat_history)\n",
" tool_calls = llm.get_tool_calls_from_response(\n",
" resp, error_on_no_tool_call=False\n",
" )\n",
" print(\"Final response: \", resp.message.content)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"We can also call multiple tools simultaneously in a single request, making it efficient for complex queries that require different types of information."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Model made 2 tool calls:\n",
"1. get_current_time with args: {'timezone': 'America/New_York'}\n",
"2. get_temperature with args: {'city': 'New York'}\n"
]
}
],
"source": [
"# Define another tool for temperature\n",
"def get_temperature(city: str) -> dict:\n",
" \"\"\"Get the current temperature for a city\"\"\"\n",
" return {\n",
" \"city\": city,\n",
" \"temperature\": \"25°C\",\n",
" }\n",
"\n",
"\n",
"# Create tools from functions\n",
"tool1 = FunctionTool.from_defaults(fn=get_current_time)\n",
"tool2 = FunctionTool.from_defaults(fn=get_temperature)\n",
"\n",
"# Ask a question that requires both tools\n",
"chat_history = [\n",
" ChatMessage(\n",
" role=\"user\",\n",
" content=\"What is the current time and temperature in New York?\",\n",
" )\n",
"]\n",
"\n",
"# The model will intelligently decide which tools to call\n",
"resp = llm.chat_with_tools([tool1, tool2], chat_history=chat_history)\n",
"tool_calls = llm.get_tool_calls_from_response(\n",
" resp, error_on_no_tool_call=False\n",
")\n",
"\n",
"print(f\"Model made {len(tool_calls)} tool calls:\")\n",
"for i, tool_call in enumerate(tool_calls, 1):\n",
" print(f\"{i}. {tool_call.tool_name} with args: {tool_call.tool_kwargs}\")"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"## Google Search Grounding\n",
"\n",
"Google Gemini 2.0 and 2.5 models support Google Search grounding, which allows the model to search for real-time information and ground its responses with web search results. This is particularly useful for getting up-to-date information.\n",
"\n",
"The `built_in_tool` parameter accepts Google Search tools that enable the model to ground its responses with real-world data from Google Search results."
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"The next total solar eclipse visible in the United States will occur on August 23, 2044. However, totality will only be visible in Montana, North Dakota, and South Dakota. Another total solar eclipse will occur on August 12, 2045, with a path spanning from California to Florida.\n",
"\n"
]
}
],
"source": [
"from llama_index.llms.google_genai import GoogleGenAI\n",
"from llama_index.core.llms import ChatMessage\n",
"from google.genai import types\n",
"\n",
"# Create Google Search grounding tool\n",
"grounding_tool = types.Tool(google_search=types.GoogleSearch())\n",
"\n",
"llm = GoogleGenAI(\n",
" model=\"gemini-2.5-flash\",\n",
" built_in_tool=grounding_tool,\n",
")\n",
"\n",
"resp = llm.complete(\"When is the next total solar eclipse in the US?\")\n",
"print(resp)"
]
},
{
"cell_type": "markdown",
"metadata": {},
"source": [
"The Google Search grounding tool provides several benefits:\n",
"\n",
"- **Real-time information**: Access to current events and up-to-date data\n",
"- **Factual accuracy**: Responses grounded in actual search results\n",
"- **Source attribution**: Grounding metadata includes search sources\n",
"- **Automatic search decisions**: The model determines when to search based on the query\n",
"\n",
"You can also use the grounding tool with chat messages:"
]
},
{
"cell_type": "code",
"execution_count": null,
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"assistant: Spain won Euro 2024, defeating England 2-1 in the final. The match took place at the Olympiastadion in Berlin. This victory marks Spain's fourth European Championship title, surpassing Germany for the most wins in the competition.\n",
"\n",
"{'grounding_chunks': [{'retrieved_context': None, 'web': {'domain': None, 'title': 'olympics.com', 'uri': 'https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQEkqnG_iRjkf89rilwO5fSBjbAADgm-Ad83fhYOhtAgW2qoG5Y8Gkselc-GshmvpqgMzke0vSUmkc6B8WwmXuxGBl9IPk3YWsytW2nOvGo1n8MlxqcrCpP62vvqjYFoo3wDQsb-tZ3RfZYTjKSTdKfVEBhvSfi4wSKMIgbnQkRx50DLqr2w3sjYI3hyZGWdsFyJFfviXdPSnVCZqQ=='}}, {'retrieved_context': None, 'web': {'domain': None, 'title': 'aljazeera.com', 'uri': 'https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQFHwRYxryu8EgG5hG-Gwgdn9sRn88H8iehIOG7KPis7rpJcRo35EAc0onyC_5hqcjUozIddtikyjHmUdK2oIBX8_3ENpLTqpu8TyYb97EibGX6_-ZtRtlPnOsd4TukiRVwfiWMk5sk9FZCsNUEFTWb9OJzPhSjOiAPW78aoAQkM9LSKLBY5vBNyQtUsNvb7k6WEd23pHAKtofxi5i7W_qYrtZPiSkqOBTqtyJ2N69oYDw=='}}, {'retrieved_context': None, 'web': {'domain': None, 'title': 'wikipedia.org', 'uri': 'https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQF2WEgQILX6A9y0uLZzBXY9UsduYELn9ahnW-FBNNHBvTQPWkuc_9cwyKmUEbfx0iton_BcIGh_85ibG5hkoE3kPvyBFfh6dEdy3UG2Vvn9gIprxruYLiUKtx8o6I06ZyFiERJqUzboU8s8Dvbd'}}, {'retrieved_context': None, 'web': {'domain': None, 'title': 'thehindu.com', 'uri': 'https://vertexaisearch.cloud.google.com/grounding-api-redirect/AUZIYQHqXK-zKOuGkYtQFyc48K49_TYwib-bRIvPqnn5UmjUcVI69vTxIiXnpXXkJtSMHa5-cBZ6Ht_4cAuWs5GuKZSzHeAQ-sHJQ2BEk52qIzjTvSteXGf7v0oBOQ_AUTqdTOpH8vXEVhqnp3o6WFVchKfexDT2sk1IDBqlqLxqQrKD9PrMsMOvU8_kfuGqH3IR_V2GHHnrPgwgR93LpiYvFdtVDlo3Wi12kj1FAgqDHHjkqyZpSc-pJ-522x0VgcdKGX6mXZ0Ssd7-aLK0YYO028ex6-o8ZeKEqeSpC9H7GP3bnw=='}}], 'grounding_supports': [{'confidence_scores': [0.97524184, 0.950235, 0.64699775], 'grounding_chunk_indices': [0, 1, 2], 'segment': {'end_index': 55, 'part_index': None, 'start_index': None, 'text': 'Spain won Euro 2024, defeating England 2-1 in the final'}}, {'confidence_scores': [0.9290034, 0.9209086], 'grounding_chunk_indices': [2, 3], 'segment': {'end_index': 109, 'part_index': None, 'start_index': 57, 'text': 'The match took place at the Olympiastadion in Berlin'}}, {'confidence_scores': [0.842964, 0.0068578157], 'grounding_chunk_indices': [2, 1], 'segment': {'end_index': 229, 'part_index': None, 'start_index': 111, 'text': \"This victory marks Spain's fourth European Championship title, surpassing Germany for the most wins in the competition\"}}], 'retrieval_metadata': {'google_search_dynamic_retrieval_score': None}, 'retrieval_queries': None, 'search_entry_point': {'rendered_content': '\\n