📊 View Lecture Slides Full-screen presentation with navigation

Session 8: Grounding AI — Embeddings Basics

Session Duration: 2 Hours     Block: 2 — AI-Assisted Engineering & Integration

Note on Stack as of August 2026: We use the embedding model gemini-embedding-2 via @google/genai. Confirm on ai.google.dev if the recommended embedding model ID changes.

Session clock

Minutes Mode Focus
0–50 Lecture Theoretical Foundation & Concepts
50–110 Core lab Generate Vectors & Compute Similarity
110–120 Checkpoint Pair share / show artifact

Note: Core lab requires generating embeddings for your track seed vault (or Obsidian vault) to confirm cosine similarity works. Stretch work involves heading-based chunking and overlap.

Before you start (2 minutes)

Confirm you are on your track (not a solution tag) and that your API key is loaded:

cd ai-product-engineering
git fetch --tags
git checkout track-NN          # e.g. track-01 … track-10
cp .env.example .env           # if you have not already
# paste GEMINI_API_KEY=… into .env (from Session 1 / AI Studio)
npm install

Track seeds ship ~7–8 Markdown notes under sample-vault/. That is enough for today’s Core. Grow toward ≥10 notes before showcase.


Learning Objectives

By the end of this session, students will be able to:

  • Explain what vector embeddings are and how they mathematically represent semantic meaning.
  • Generate embeddings from text using the Gemini API in a Node.js environment.
  • Calculate cosine similarity in JavaScript to determine how closely related two pieces of text are.
  • Describe how embeddings form the retrieval foundation of a Retrieval-Augmented Generation (RAG) system.

Part 1: Theoretical Foundation — How AI Represents Meaning

1.1 The Problem with Static Knowledge

Every large language model has a knowledge cutoff date. A model trained in 2023 knows nothing about the election results of 2024, the internal API docs you wrote yesterday, or the specific product prices for your startup.

If you ask an ungrounded model about these topics, it will either refuse to answer or, worse, hallucinate a plausible-sounding lie.

Grounding is the process of feeding the model relevant, factual information at query time. We do this by retrieving the right document from your exocortex (your Obsidian vault) and inserting it into the prompt. But how do we find the “right” document out of thousands of notes when a user types a messy, conversational question? Keyword search (like CTRL+F) fails if the user searches for “canines” but the document says “dogs.”

The solution is semantic search powered by embeddings.

1.2 What Is a Vector Embedding?

To a computer, text is just a sequence of ASCII or UTF-8 characters. It has no meaning. AI researchers solved this by converting concepts into geometry.

A vector embedding is a long list of numbers (an array of floating-point values) that represents the semantic meaning of a piece of text. The core rule of embeddings: Similar meaning → similar numbers (closer together in space).

Imagine a simple 3D graph where the X-axis is “Animals”, the Y-axis is “Math”, and the Z-axis is “Food”.

  • The word “Puppy” might have coordinates [0.9, 0.01, 0.2].
  • The word “Dog” might be [0.85, 0.05, 0.15].
  • The word “Algebra” might be [0.01, 0.95, 0.0].

In this space, “Puppy” and “Dog” are geometrically close to each other. “Algebra” is very far away. Real embedding models (like gemini-embedding-2) don’t use 3 dimensions; they use hundreds or thousands of dimensions (commonly 768, 1536, or 3072) to capture nuanced semantic relationships, tone, and context.

1.3 Cosine Similarity: Measuring Distance

Once we convert sentences into vectors (arrays of numbers), how do we determine which ones are closest? We use a mathematical formula called Cosine Similarity. It measures the angle between two vectors.

In NLP (Natural Language Processing), cosine similarity yields a score between -1 and 1 (though practically, it’s usually between 0.0 and 1.0 when dealing with text embeddings).

  • 1.0: Identical meaning (or the exact same text).
  • 0.7–0.9: Very highly similar or highly relevant.
  • 0.5–0.7: Somewhat related (sharing some semantic overlap).
  • 0.0–0.3: Completely unrelated topics.

At query time, the math is simple:

  1. Embed the user’s question into a vector.
  2. Calculate the cosine similarity score against every chunk in your database.
  3. Keep the top K chunks (e.g., the 3 chunks with the highest scores).

1.4 The Embedding Workflow

Building a semantic search system is a two-step process.

Phase 1: Offline Indexing (Done once, or on a cron job) This happens when you build your knowledge base, before the user ever asks a question.

  1. Read the Markdown files from your vault.
  2. Split them into chunks (our Core script embeds up to 10 files, truncated per file).
  3. Call the API (embedContent) for each chunk to get its vector.
  4. Store the text chunk and its vector array together in a database (for this course, a simple embeddings.json file — gitignored).

Phase 2: Online Retrieval (Done at query time — Session 9) This happens instantly when a user clicks ‘Submit’.

  1. Take the user’s text query.
  2. Call the API to embed the query.
  3. Run cosine similarity to find the nearest stored vectors.
  4. Inject those matching text chunks into the Gemini prompt.

1.5 Chunking Strategy

You cannot embed an entire 50-page book as a single vector. The “meaning” gets diluted. You must chunk the text.

  • Target Chunk Size: 200–500 words is a sweet spot for RAG. It’s enough to provide context, but small enough to remain semantically focused.
  • Semantic Boundaries: Do not just split text arbitrarily every 1,000 characters (which might cut a sentence in half). Split on natural boundaries like Markdown headings (##) or double line breaks (paragraphs).
  • Overlap: (Advanced / Stretch) When chunking, include a 50–100 word overlap between chunk A and chunk B so ideas that cross a boundary are not lost.

1.6 Important API note for gemini-embedding-2

With the older model gemini-embedding-001, you could pass config: { taskType: 'RETRIEVAL_DOCUMENT' }.

For gemini-embedding-2, taskType is not supported. Google’s docs ask you to put the task in the text prefix instead (asymmetric retrieval):

Role Format the string as
Document (indexing) title: none \| text: {content} (or use a real title)
Query (search) task: search result \| query: {content}

Our embedText(ai, text, taskType) helper still accepts a third argument for compatibility with the kit script. Internally you will map RETRIEVAL_DOCUMENT / RETRIEVAL_QUERY onto those prefixes — do not send taskType in config.


Part 2: Practical Labs — Generate and Explore Embeddings

Lab 8.1 — Implement embedText (Core)

Open lib/embeddings.js. cosineSimilarity is already done. Implement embedText so it returns a number[].

The build script (scripts/build-embeddings.js) already creates GoogleGenAI and passes it in — use the ai argument; do not create a second client inside this file.

const EMBEDDING_MODEL = 'gemini-embedding-2';

export async function embedText(ai, text, taskType = 'RETRIEVAL_DOCUMENT') {
  const raw = String(text ?? '');
  // gemini-embedding-2: task goes in the text, not config.taskType
  const contents =
    taskType === 'RETRIEVAL_QUERY'
      ? `task: search result | query: ${raw}`
      : `title: none | text: ${raw}`;

  const response = await ai.models.embedContent({
    model: EMBEDDING_MODEL,
    contents,
  });

  const values = response.embeddings?.[0]?.values;
  if (!values?.length) {
    throw new Error('embedContent returned no embedding values');
  }
  return values;
}

Smoke-test (from the project root, with .env set):

node --input-type=module -e "import 'dotenv/config'; import { GoogleGenAI } from '@google/genai'; import { embedText } from './lib/embeddings.js'; const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY }); const v = await embedText(ai, 'I love coffee'); console.log('dims', v.length, 'sample', v.slice(0,5));"

You should see a positive dims count (often 3072, or fewer if the API returns a reduced size). If you get Set GEMINI_API_KEY / 401 / 429, fix the key or wait and retry — do not invent fake vectors.

Lab 8.2 — Compute Cosine Similarity (Core)

cosineSimilarity(vecA, vecB) is already in lib/embeddings.js. Prove semantic search with three strings:

node --input-type=module -e "import 'dotenv/config'; import { GoogleGenAI } from '@google/genai'; import { embedText, cosineSimilarity } from './lib/embeddings.js'; const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY }); const A = await embedText(ai, 'The artificial intelligence model is training.'); const B = await embedText(ai, 'Machine learning algorithms are optimizing.'); const C = await embedText(ai, 'I prefer my espresso with oat milk.'); console.log('A~B', cosineSimilarity(A,B).toFixed(3)); console.log('A~C', cosineSimilarity(A,C).toFixed(3));"

Observation: A and B use different words but should score higher than A and C. That is semantic similarity, not keyword match.

Optional polish: For a pure similarity drill (not retrieval), Google recommends the symmetric prefix task: sentence similarity | query: … on all three strings. The Core embedText helper uses retrieval prefixes — that is correct for Lab 8.3 / Session 9; relative A~B vs A~C still works for the classroom demo. (All three calls use the same default document formatting so the scores are comparable; Session 9 will embed queries with the RETRIEVAL_QUERY prefix.)

Lab 8.3 — Embed the Vault Notes (Core)

Run Phase 1 (offline indexing) with the kit script. It walks Markdown files, calls your embedText, and writes embeddings.json in the project root (gitignored).

  1. Ensure Lab 8.1 is implemented and exported.
  2. Prefer your track’s sample-vault/ for Core (default path). Cap is 10 files.
# Default: ./sample-vault (works on track branches)
npm run build-embeddings

Own Obsidian vault?

# bash / macOS / Git Bash
VAULT_PATH="/path/to/your/obsidian/vault" npm run build-embeddings

# Windows PowerShell
$env:VAULT_PATH="C:\path\to\your\obsidian\vault"; npm run build-embeddings
  1. Open embeddings.json. You should see an array of objects shaped like { file, content, embedding } — the vector array key is embedding, not vector.
  2. Quick length check:
node --input-type=module -e "import data from './embeddings.json' with { type: 'json' }; console.log('notes', data.length, 'dims', data[0]?.embedding?.length);"

Do not proceed to Session 9 until embeddings.json exists and has at least a few entries with non-empty embedding arrays.

Recovery: If the script says Did you implement embedText()?, finish Lab 8.1. If No .md files, you are probably on main without a vault — git checkout track-NN. If the API fails, keep screenshots of the error and retry later (see GATES.md).

Stretch Goal: Modify scripts/build-embeddings.js to split on ## headings (and optionally add overlap) instead of embedding each file as one truncated blob.


Key Takeaways

  • Vector Embeddings translate semantic meaning into geometry (lists of numbers). Sentences with similar meanings produce similar vectors, regardless of the exact vocabulary used.
  • Cosine similarity measures relevance between a query vector and stored document vectors.
  • RAG needs an offline indexing phase (embeddings.json) and an online retrieval phase (Session 9).
  • For gemini-embedding-2, put retrieval task hints in the text prefix, not in config.taskType.
  • You now have a mathematical index of your exocortex, ready to search in Session 9.

Further Reading & Resources

  • Gemini Embeddings Guide: ai.google.dev/gemini-api/docs/embeddings — model IDs, prefixes for Embedding 2, dimensions.
  • “What are Word Embeddings?” — conceptual explainers by Jay Alammar or 3Blue1Brown (YouTube) for high-dimensional intuition.
  • Classroom unstick: HOW-TO-USE-SOLUTIONS.md. Caution: published solution-NN-phase-3 tags may still pass config: { taskType }, which Google does not support on gemini-embedding-2. Prefer the Lab 8.1 snippet on this page; if you peek at a solution, rewrite it to use text prefixes before you copy it onto your track.