Agentic AI โ Day 3 (After Hours)
Day 3 Extra Content โ Building the RAG Pipeline in Flowise
What We're Going to Cover
In the required Day 3 exercises, you grounded FitPath's workout-plan prompt in NotebookLM and checked whether the grounding held up. That worked โ but NotebookLM made every architectural decision for you. It chose how to chunk your documents, which embedding model to use, how many chunks to retrieve. You never saw any of it.
This is where you build the thing yourself. Same FitPath documents, same question โ but this time you assemble the pipeline node by node: document loader, text splitter, embedding model, vector store, retriever, LLM. Every knob from today's material becomes something you actually set, not something a product decided for you.
By the end you will have:
- Built a working RAG pipeline from scratch and watched it answer a question using your own retrieved chunks
- Changed chunk size and retrieval count and watched the actual retrieved chunks change โ not just the final answer
- Compared what a transparent pipeline exposes that a black-box tool like NotebookLM doesn't
Overview
- Time: ~75โ90 minutes total, across three exercises
- Difficulty: Moderate-to-high โ first-time tool setup takes longer than the exercise itself
- Format: Individual, after hours
- Setup: Your own Flowise Cloud account (free tier) + your own OpenAI API key. See below.
Learning Objective
Assemble a RAG pipeline's components explicitly (chunking, embedding, retrieval) and observe how changing chunk size and top-k retrieval count changes what the model actually sees before it answers.
Relevant Day 3 Material
How Vector Search Works, Data Ingestion, Chunking / Chunking Tradeoffs, Basic RAG Flow, Retrieval Quality Is Crucial.
Setup โ Read Before You Start
This setup takes 15โ20 minutes the first time. Budget for it separately from the exercise time above.
- Create your own Flowise Cloud account at flowise.ai โ sign up individually, don't share an account with a classmate. The free tier gives you 100 predictions/month and 2 flows, which is enough for this exercise, but only if it's yours alone.
- Get an OpenAI API key at platform.openai.com/api-keys. You'll need to add a payment method, but the entire exercise โ embeddings plus a handful of chat completions โ will cost a few cents, not dollars.
- Reuse the three FitPath source documents from the required exercises:
FitPath_Source_1_Exercise_Library_and_Programming_Guide.mdFitPath_Source_2_Safety_and_Injury_Guidelines_CURRENT.mdFitPath_Source_3_Safety_and_Injury_Guidelines_OUTDATED.md(not needed until you're done โ set it aside)
- Use the same user profile from the required exercises: 29-year-old, first week on FitPath, gym access, mild recurring knee pain, no medical clearance on file.
Agentic AI โ Day 3 (After Hours)
Exercise 1 โ Build the Pipeline
What You're Building
A minimal but real RAG pipeline: Text Splitter โ File Loader โ Embeddings โ In-Memory Vector Store โ Chat Model โ Conversational Retrieval QA Chain. You don't need a dedicated vector database for this โ Flowise's in-memory vector store is enough for a document set this size and avoids a third signup. There's no separate "Retriever" node in Flowise โ retrieval settings (top-k) live directly on the Vector Store node.
โ ๏ธ General rule for the rest of this exercise: every time you change a node's settings โ a prompt, a chunk size, a top-k value โ click Save on the canvas (the main flow-level save, not just closing the node panel) before you re-run the chat. Flowise does not reliably apply node edits to the running chain until the whole flow is saved again. If you change something and nothing seems to happen, this is almost always why.
How to Run It
- Add a Recursive Character Text Splitter node. Set Chunk Size to 500 and Chunk Overlap to 50. It won't connect to anything yet โ you'll wire it into the loader next.
- Add a File Loader node. Connect the Text Splitter's output into the File Loader's Text Splitter input slot. Upload
FitPath_Source_1_Exercise_Library_and_Programming_Guide.mdandFitPath_Source_2_Safety_and_Injury_Guidelines_CURRENT.mddirectly on this node โ one File Loader handles multiple files. - Add an OpenAI Embedding node. Connect your OpenAI credential and select an embedding model (
text-embedding-ada-002works fine for this exercise). - Add an In-Memory Vector Store node. Make two connections into it: File Loader's Document output โ Vector Store's Document input, and OpenAI Embedding's output โ Vector Store's Embeddings input. Set Top K to 4 in this node's own settings. On the output dropdown at the bottom, select Memory Retriever โ this is what exposes it as a retriever the QA Chain can use.
- Add a ChatOpenAI node. Connect your OpenAI credential and pick a chat model (
gpt-4o-miniis a good default). Set Temperature to around 0.2โ0.3 โ keep it low so that when your answers change in Exercise 2, you know it's because retrieval changed, not random variance. - Add a Conversational Retrieval QA Chain node. Connect ChatOpenAI's output โ Chat Model input, and the Vector Store's Memory Retriever output โ Vector Store Retriever input. Leave Memory and Input Moderation empty. Turn Return Source Documents ON โ you need to see the actual retrieved chunks, not just the final answer.
- Fix the default Response Prompt before you test. Open Additional Parameters on the QA Chain node. Flowise's default Response Prompt instructs the model to say "Hmm, I'm not sure" whenever it can't find a literal, pre-written answer sitting in the retrieved text โ and that's exactly what will happen here, because your source documents are rules to apply, not a pre-written plan to quote back. Left as default, this pipeline will refuse to generate anything. Replace the Response Prompt with:
You are FitPath's workout plan assistant. Use the following context to generate a personalised weekly workout plan for the user, applying the exercise library and safety rules given. Compose the plan yourself based on these rules โ this is not a lookup task. Only decline if the user's request falls outside what the safety rules cover. {context} Question: {question} Answer in Markdown:
{context}and{question}must be typed exactly like that, curly braces included โ Flowise substitutes them automatically at runtime. Leave the Rephrase Prompt as its default; it already works correctly. - Save the whole chatflow โ the main Save button on the canvas, not just the node panel. This is the step it's easiest to forget, and skipping it is the single most common reason your Step 7 edit won't seem to do anything.
- Open the chat panel and ask the same question you asked NotebookLM in the required exercises:
Generate a personalised weekly workout plan for the user below. Output as a day-by-day schedule with exercises, sets, and reps. User: 29 years old. Goal: general fitness. First week ever using FitPath โ complete beginner. Trains 3 days/week. Has access to a full commercial gym (Gym Standard). Reports mild, recurring knee pain from an old injury โ no current medical clearance on file. No other conditions reported.If you get "Hmm, I'm not sure" as the answer, it means either the Response Prompt edit from Step 7 didn't include
{context}and{question}, or the flow wasn't re-saved after the edit (Step 8) โ go back and check both before re-running.
Paste your full pipeline's answer here:
Flowise (unlike NotebookLM) can show you the raw retrieved chunks, not just the final answer โ look for a "view sources" or similar option on the response. List the actual chunks that were retrieved, not just which document they came from:
Compare this answer to your NotebookLM grounded answer from the required exercises. Same conclusions, or different? If different, is that a chunking difference, a retrieval-count difference, or something else?
Look at the retrieved chunks you listed. Did the retriever pull in the right section of Source 2 for the knee-pain rule (Section 2), or did it also โ or instead โ retrieve something less relevant, like the pregnancy section?
Success Criteria
- A working pipeline that produces an answer without errors
- The actual retrieved chunks recorded, not just the final answer
- A specific comparison against your NotebookLM result โ not "similar" but naming what matched or didn't
Agentic AI โ Day 3 (After Hours)
Exercise 2 โ The Chunk Size Experiment
What You're Testing
Earlier today you saw the small-chunk-vs-large-chunk tradeoff as a table. Now you get to run both settings against the same question and see which one actually happens.
How to Run It
- Edit the Text Splitter node directly on your existing Exercise 1 flow โ don't duplicate the flow. Duplicating a chatflow in Flowise Cloud has a known issue where uploaded files don't carry over correctly to the copy, producing a
NoSuchKeyerror when you try to run it. Editing your working flow in place avoids this entirely. - Set chunk size to 150 characters, overlap 20. Save the whole chatflow, then re-run the same workout-plan question. Record the full answer and the full list of retrieved chunks below.
- Reset to chunk size 1500 characters, overlap 100. Save the flow again, then re-run the same question. Record the full answer and the full list of retrieved chunks below.
- Your 500/50 result from Exercise 1 is your "original" baseline โ no need to rerun it. Just carry that answer and those chunks forward into the comparison table.
Record What You Get โ Small Chunk Size (150 Chars, Overlap 20)
Paste the full answer here:
List the actual retrieved chunks:
Record What You Get โ Large Chunk Size (1500 Chars, Overlap 100)
Paste the full answer here:
List the actual retrieved chunks:
Comparison Table
Fill this in using the full answers and chunks you just recorded, plus your Exercise 1 result for the "Original" row.
| Setting | Retrieved chunks (summary) | Answer changed? How? |
|---|---|---|
| Small (150 chars) | ||
| Original (500 chars) โ from Exercise 1 | ||
| Large (1500 chars) |
At the small chunk size, did retrieval fragment a rule across multiple chunks in a way that lost meaning โ for example, splitting the knee-pain rule from its exception or its reasoning?
At the large chunk size, did the retriever pull in a chunk so broad it included irrelevant rules alongside the relevant one (e.g., knee guidance bundled with unrelated shoulder guidance)?
Based on what you saw, what chunk size would you actually recommend for Source 2 in production โ and why? Tie your answer to the specific failure mode you observed, not the general tradeoff covered earlier.
Success Criteria
All three rows filled with genuine differences observed (not assumed), and a chunk-size recommendation grounded in something you actually saw fail at the other two settings.
Agentic AI โ Day 3 (After Hours)
Exercise 3 โ Retrieval Quality Stress Test
What You're Testing
Retrieval quality is the biggest determinant of RAG quality โ and RAG evaluation must test retrieval, not just final answers. You now have a tool that shows you retrieval directly. Use it.
How to Run It
Reset your Text Splitter back to your Exercise 1 settings (500/50) and save the flow. Run each of these three tests, and for each one, check the actual retrieved chunks, not just whether the final answer sounded reasonable. Remember to save the flow after each settings change before re-running.
Test A โ Right question, wrong top-k. Set the Vector Store node's top-k down to 1. Ask: "What's the session length rule for a beginner, and what should I know about this user's knee?" โ a question that needs information from two different places in Source 2. Does one chunk retrieval starve the answer of half of what it needs?
retrieved chunks:
full answer:
what this tells you:
Test B โ The boundary question. Ask something outside both documents entirely: "Should this user also change their diet to support the training plan?" Does your pipeline correctly have nothing relevant to retrieve โ and if so, does the LLM say so, or does it answer anyway from general knowledge? Note: your Response Prompt from Exercise 1 already instructs the model to "only decline if the request falls outside what the safety rules cover" โ so this test is checking whether that instruction actually holds, not testing Flowise's untouched default behavior.
retrieved chunks (if any):
full answer:
what this tells you:
Test C โ Your own adversarial question. Write one question designed to expose a retrieval weakness you'd expect based on Exercise 2's findings. State your hypothesis before you run it.
hypothesis:
retrieved chunks:
full answer:
did the result match your hypothesis?:
In Test A, was low top-k a real problem for this question, or did one chunk happen to be enough? What does that tell you about setting top-k for a real feature, where you can't hand-pick easy questions?
In Test B โ compare this to what happened when you asked NotebookLM an out-of-scope question in the required exercises. Same behavior, or different? If different, what's the architectural reason โ is it the retrieval design, the LLM's own instructions, or something you didn't set at all?
Success Criteria
All three tests run with actual retrieved-chunk evidence (not inferred from the answer alone), Test C's hypothesis stated before running it, and a direct comparison to the NotebookLM boundary-question result from the required exercises.
Closing the Loop
You've now seen the same FitPath grounding problem solved three ways: no grounding at all (Day 2's ungrounded prompt), a black-box grounded tool (NotebookLM), and a pipeline where you set every parameter yourself (Flowise). None of the three is automatically "best" โ that's exactly the decision the RAG Readiness Review and the Prompt-Only vs. RAG vs. Fine-Tuning framework are for. Keep your Flowise flow and your NotebookLM notebook both intact โ you may need to compare them again.
