Project Track 2: Smart Pantry & Recipe Architect

The Problem / Concept Deciding what to cook based on fragmented leftover ingredients and specific household dietary restrictions.

Project Overview & Objectives

Food waste is a major issue, and meal planning is tedious when balancing what’s expiring in the fridge with complex household dietary restrictions. This application acts as a personal chef that generates recipes dynamically, ensuring no allergens are included and utilizing available ingredients.

If you select this track, your goal is to build a functional Minimum Viable Prototype (MVP) that seamlessly integrates a Node.js/Express backend with Google’s Gemini API, utilizing Retrieval-Augmented Generation (RAG), multimodal vision, and autonomous tool calling.


Detailed Requirements Document (PRD)

1. RAG (Obsidian) Core Requirement

To prevent hallucination, the AI must be grounded in a specific, personal knowledge base. You will build this using Markdown files in Obsidian.

  • Knowledge Base Content: The knowledge base consists of Markdown files detailing household profiles (e.g., ‘Dad is lactose intolerant’, ‘Kid allergic to peanuts’) and a database of trusted family recipes. The RAG pipeline must retrieve the dietary restrictions before generation to ensure safety.
  • Implementation Expectation: Your Express server must parse these Markdown files, generate vector embeddings using gemini-embedding-2, and perform a cosine similarity search against the user’s query. The retrieved text must be injected into the system prompt.

2. Multimodal (Vision) Stretch Goal

AI is not just text. Modern products must perceive the world.

  • Vision Use Case: Users can upload a photo of an open fridge or pantry shelf. The system will prompt Gemini to identify visible ingredients, return them as a structured list, and then cross-reference them with the dietary restriction RAG context to propose a safe, viable meal.
  • Implementation Expectation: Your frontend HTML must include a file upload input. The image must be converted to base64, sent to the /query endpoint, and passed to gemini-2.5-flash alongside the text prompt and RAG context.

3. Tool Calling Stretch Goal

Agents need to take actions in the real world or fetch real-time data that isn’t in their RAG database.

  • Tool Definition: Implement a calculate_nutrition(ingredients) tool. When the AI finalizes a recipe, it calls this tool (which simulates hitting the Edamam or USDA API) to fetch calorie counts and macronutrient breakdowns, displaying a final nutritional summary to the user.
  • Implementation Expectation: You must define a strict JSON schema for this tool and register it in your Gemini API call. When the model decides to invoke the tool, your Node.js server must intercept the request, execute a mock function, and return the result to the model for the final response.

Step-by-Step Implementation Guide

If you are using this document to prompt an AI coding assistant (like GitHub Copilot or Cursor), use the following phasing:

  1. Phase 1: UI & Mock Backend: Ask the AI to generate a vanilla HTML/JS interface with a text input, file upload, and a submit button. Connect it to an Express POST /query route that returns mock JSON.
  2. Phase 2: Vanilla Gemini Integration: Connect the Express route to the actual @google/genai SDK. Send a simple text prompt and display the result.
  3. Phase 3: The RAG Pipeline: Ask the AI to write a script to read your Markdown files, split them into chunks, and get embeddings. Update your /query route to calculate cosine similarity and inject the top matches.
  4. Phase 4: Multimodal & Tools: Finally, add the base64 image parsing to the API call, and define your function declaration for the tool-calling stretch goal.

Note: In the future, a specific branch in the course repository will be provided containing starter scaffolding tailored to this exact project track.