Project Track 3: Local Hardware Troubleshooting Bot
The Problem / Concept Non-technical users need help fixing specific household or office equipment without sorting through generic web forums.
Project Overview & Objectives
When the internet goes down or a printer breaks, generic AI advice (‘Try restarting your router’) is frustrating. This bot knows the exact model numbers, IP addresses, and idiosyncrasies of your specific local hardware setup, acting as an instant IT support technician.
If you select this track, your goal is to build a functional Minimum Viable Prototype (MVP) that seamlessly integrates a Node.js/Express backend with Google’s Gemini API, utilizing Retrieval-Augmented Generation (RAG), multimodal vision, and autonomous tool calling.
Detailed Requirements Document (PRD)
1. RAG (Obsidian) Core Requirement
To prevent hallucination, the AI must be grounded in a specific, personal knowledge base. You will build this using Markdown files in Obsidian.
- Knowledge Base Content: Ingest Markdown versions of PDF manuals for specific devices (e.g., ‘Epson L3150 Printer’, ‘Netgear Nighthawk Router’), alongside a ‘Network_Topology.md’ file that lists static IP addresses and admin passwords. The bot must look up specific error codes from these manuals.
- Implementation Expectation: Your Express server must parse these Markdown files, generate vector embeddings using
gemini-embedding-2, and perform a cosine similarity search against the user’s query. The retrieved text must be injected into the system prompt.
2. Multimodal (Vision) Stretch Goal
AI is not just text. Modern products must perceive the world.
- Vision Use Case: Allow users to upload photos of blinking LED error sequences on a router or an obscure error screen on a smart TV. The AI translates the visual state into an error code and retrieves the exact mitigation steps from the manuals.
- Implementation Expectation: Your frontend HTML must include a file upload input. The image must be converted to base64, sent to the
/queryendpoint, and passed togemini-2.5-flashalongside the text prompt and RAG context.
3. Tool Calling Stretch Goal
Agents need to take actions in the real world or fetch real-time data that isn’t in their RAG database.
- Tool Definition: Create a
ping_device(ip_address)orcheck_internet_status()tool. The AI can decide to use this tool to determine if the issue is a local hardware failure (cannot ping printer) or an ISP outage (cannot ping Google), altering its troubleshooting advice accordingly. - Implementation Expectation: You must define a strict JSON schema for this tool and register it in your Gemini API call. When the model decides to invoke the tool, your Node.js server must intercept the request, execute a mock function, and return the result to the model for the final response.
Step-by-Step Implementation Guide
If you are using this document to prompt an AI coding assistant (like GitHub Copilot or Cursor), use the following phasing:
- Phase 1: UI & Mock Backend: Ask the AI to generate a vanilla HTML/JS interface with a text input, file upload, and a submit button. Connect it to an Express
POST /queryroute that returns mock JSON. - Phase 2: Vanilla Gemini Integration: Connect the Express route to the actual
@google/genaiSDK. Send a simple text prompt and display the result. - Phase 3: The RAG Pipeline: Ask the AI to write a script to read your Markdown files, split them into chunks, and get embeddings. Update your
/queryroute to calculate cosine similarity and inject the top matches. - Phase 4: Multimodal & Tools: Finally, add the base64 image parsing to the API call, and define your function declaration for the tool-calling stretch goal.
Note: In the future, a specific branch in the course repository will be provided containing starter scaffolding tailored to this exact project track.


