AI Product Engineering
Block 2: AI-Assisted Engineering & Integration
Model knows nothing after its training date
Model knows nothing about your domain, your documents, your data
Grounding = feed relevant information at query time, not from training data.
A list of numbers that represents the meaning of text.
Key property: similar meaning → similar numbers.
Real models: hundreds to thousands of dimensions.
| Score | Meaning |
|---|---|
| 1.0 | Identical meaning |
| 0.7–0.9 | Very similar |
| 0.5–0.7 | Related but different |
| 0.0 | Completely unrelated |
Find relevant documents = find highest similarity to query embedding.
Which pair should score highest on cosine similarity?
Choose the best definition.
Choose the best interpretation of the score.
Read each document → call embedding API → store text + vector
Embed query → find most similar vectors → feed matching text to LLM
True or false.
True or false.
True or false.
gemini-embedding-2config.taskType — not supported on Embedding 2title: none | text: …task: search result | query: …response.embeddings[0].values as a number[]git checkout track-NN · GEMINI_API_KEY in .env · npm installembedText in lib/embeddings.js (prefix format, no taskType)npm run build-embeddings (default sample-vault/, max 10 files)embeddings.json has { file, content, embedding } — Session 9 gates on thisStretch: heading split + overlap chunking. Seeds are ~7–8 notes; grow to ≥10 before showcase.
config.taskTypeembeddings.json ready — Session 9 builds RAG on topNext Session: Building the RAG Pipeline