Wall of love for The Gen Academy
I spent 3 months building a multi-agent system that does 2-4 hours of AE work in ~4 minutes.
4 LangGraph agents. Cited sales strategy. Production-grade.
3 things production taught me:
-> NAIVE RAG FAILS HARD
"GNC toolboxes" returned zero semantic similarity to "Sensor Fusion and Tracking Toolbox." Vector-only search = wrong products in front of customers.
Fixed it with Hybrid RAG (thanks Aishwarya Srinivasan): BM25 + vector + cross-encoder reranking. Retrieval quality is the ceilingโeverything downstream inherits its failures.
-> YOU CAN'T TUNE WITHOUT EVALS
Like adjusting PID gains without feedback means you're flying blind.
Built Claude-as-judge with 4 metrics + 9 deterministic checks. Went from 3.0 to 4.2 quality score across 3 iterations.
-> LLM COST IS A ROUTING PROBLEM
Ollama ($0) โ Groq 8B โ Groq 70B dispatched by task complexity. Prompt caching cuts 90% of token costs.
MIT licensed. If any of this saves you a debugging session, fork it, build on it.
Repo Link in comments.
#AgenticAI #LangGraph #RAG
I spent 3 months building a multi-agent system that does 2-4 hours of AE work in ~4 minutes.
4 LangGraph agents. Cited sales strategy. Production-grade.
3 things production taught me:
-> NAIVE RAG FAILS HARD
"GNC toolboxes" returned zero semantic similarity to "Sensor Fusion and Tracking Toolbox." Vector-only search = wrong products in front of customers.
Fixed it with Hybrid RAG (thanks Aishwarya Srinivasan): BM25 + vector + cross-encoder reranking. Retrieval quality is the ceilingโeverything downstream inherits its failures.
-> YOU CAN'T TUNE WITHOUT EVALS
Like adjusting PID gains without feedback means you're flying blind.
Built Claude-as-judge with 4 metrics + 9 deterministic checks. Went from 3.0 to 4.2 quality score across 3 iterations.
-> LLM COST IS A ROUTING PROBLEM
Ollama ($0) โ Groq 8B โ Groq 70B dispatched by task complexity. Prompt caching cuts 90% of token costs.
MIT licensed. If any of this saves you a debugging session, fork it, build on it.
Repo Link in comments.
#AgenticAI #LangGraph #RAG
Most AI answers ๐ด๐ฐ๐ถ๐ฏ๐ฅ ๐ณ๐ช๐จ๐ฉ๐ตโฆ but are they actually correct?
One of the challenges with LLMs becomes very clear in enterprise settings.
LLMs like ChatGPT, Gemini, or Claude are trained on ๐ถ๐ป๐๐ฒ๐ฟ๐ป๐ฒ๐ ๐ฑ๐ฎ๐๐ฎ, not your companyโs internal policies, processes, or documents.
So when someone asks:
โ๐๐ข๐ฏ ๐ ๐ด๐ฉ๐ข๐ณ๐ฆ ๐ข ๐ค๐ญ๐ช๐ฆ๐ฏ๐ต ๐ง๐ช๐ญ๐ฆ ๐ธ๐ช๐ต๐ฉ ๐ข ๐ท๐ฆ๐ฏ๐ฅ๐ฐ๐ณ?โ
The response may sound confident and logicalโฆ
but it could still be ๐บ๐ถ๐๐ฎ๐น๐ถ๐ด๐ป๐ฒ๐ฑ ๐๐ถ๐๐ต ๐ฒ๐ป๐๐ฒ๐ฟ๐ฝ๐ฟ๐ถ๐๐ฒ ๐ฝ๐ผ๐น๐ถ๐ฐ๐.
๐ง๐ต๐ถ๐ ๐ถ๐ ๐๐ต๐ฒ๐ฟ๐ฒ ๐ฅ๐๐ (๐ฅ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น ๐๐๐ด๐บ๐ฒ๐ป๐๐ฒ๐ฑ ๐๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐๐ถ๐ผ๐ป) ๐ฐ๐ผ๐บ๐ฒ๐ ๐ถ๐ป
RAG connects LLMs to ๐ฒ๐ป๐๐ฒ๐ฟ๐ฝ๐ฟ๐ถ๐๐ฒ ๐ธ๐ป๐ผ๐๐น๐ฒ๐ฑ๐ด๐ฒ ๐๐ต๐ฒ๐ ๐๐ฒ๐ฟ๐ฒ ๐ป๐ฒ๐๐ฒ๐ฟ ๐๐ฟ๐ฎ๐ถ๐ป๐ฒ๐ฑ ๐ผ๐ป.
Instead of relying only on pre-trained knowledge, it:
1. Breaks documents into ๐ฐ๐ต๐๐ป๐ธ๐
2. Converts them into ๐ฒ๐บ๐ฏ๐ฒ๐ฑ๐ฑ๐ถ๐ป๐ด๐ (numerical representations of text)
3. Stores them in a ๐๐ฒ๐ฐ๐๐ผ๐ฟ ๐ฑ๐ฎ๐๐ฎ๐ฏ๐ฎ๐๐ฒ
4. Uses ๐ฐ๐ผ๐๐ถ๐ป๐ฒ ๐๐ถ๐บ๐ถ๐น๐ฎ๐ฟ๐ถ๐๐(=0.97) to retrieve the most relevant context
5. Feeds that context to the LLM to generate grounded responses
๐๐ต๐๐ป๐ธ๐ถ๐ป๐ด ๐ถ๐ ๐บ๐ผ๐ฟ๐ฒ ๐ถ๐บ๐ฝ๐ผ๐ฟ๐๐ฎ๐ป๐ ๐๐ต๐ฎ๐ป ๐ถ๐ ๐น๐ผ๐ผ๐ธ๐
The way you split data directly impacts the quality of retrieval.
Common chunking strategies include:
โข Fixed-length chunking
โข Sentence-based chunking
โข Paragraph-based chunking
โข Sliding window chunking
โข Semantic chunking
โข Recursive chunking
Better chunking โ Better retrieval โ Better answers
And more importantly:
๐๐ฒ๐๐๐ฒ๐ฟ ๐ฎ๐ป๐๐๐ฒ๐ฟ๐ ๐๐ถ๐๐ต ๐ฐ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป๐
Which means the response is not just helpful, but ๐๐ฟ๐ฎ๐ฐ๐ฒ๐ฎ๐ฏ๐น๐ฒ:
โข You can see where the answer came from
โข You can refer back to the exact company document
โข The answer has real grounding, not just confidence
Thatโs what gives AI ๐ฐ๐ฟ๐ฒ๐ฑ๐ถ๐ฏ๐ถ๐น๐ถ๐๐ ๐ถ๐ป ๐ฒ๐ป๐๐ฒ๐ฟ๐ฝ๐ฟ๐ถ๐๐ฒ ๐๐๐ฒ.
โ๏ธ๐ฅ๐๐ ๐๐ ๐๐ถ๐ป๐ฒ-๐๐๐ป๐ถ๐ป๐ด
๐๐ถ๐ป๐ฒ-๐๐๐ป๐ถ๐ป๐ด โ Good for behavior, structure, and low latency
๐ฅ๐๐ โ Better for dynamic, real-time, and private enterprise data
For most enterprise use cases, RAG is the most effective starting point for improving LLM responses.
๐ ๐ ๐๐ฎ๐ธ๐ฒ๐ฎ๐๐ฎ๐
AI is not just about generating answers.
Itโs about:
โข Retrieving the right context
โข Grounding responses in real data
โข Making answers ๐๐ฒ๐ฟ๐ถ๐ณ๐ถ๐ฎ๐ฏ๐น๐ฒ, ๐ป๐ผ๐ ๐ท๐๐๐ ๐ฐ๐ผ๐ป๐๐ถ๐ป๐ฐ๐ถ๐ป๐ด
A quick shoutout to Arvind Narayanamurthy & Aishwarya Srinivasan from The Gen Academy for this weekend learning.
Most AI answers ๐ด๐ฐ๐ถ๐ฏ๐ฅ ๐ณ๐ช๐จ๐ฉ๐ตโฆ but are they actually correct?
One of the challenges with LLMs becomes very clear in enterprise settings.
LLMs like ChatGPT, Gemini, or Claude are trained on ๐ถ๐ป๐๐ฒ๐ฟ๐ป๐ฒ๐ ๐ฑ๐ฎ๐๐ฎ, not your companyโs internal policies, processes, or documents.
So when someone asks:
โ๐๐ข๐ฏ ๐ ๐ด๐ฉ๐ข๐ณ๐ฆ ๐ข ๐ค๐ญ๐ช๐ฆ๐ฏ๐ต ๐ง๐ช๐ญ๐ฆ ๐ธ๐ช๐ต๐ฉ ๐ข ๐ท๐ฆ๐ฏ๐ฅ๐ฐ๐ณ?โ
The response may sound confident and logicalโฆ
but it could still be ๐บ๐ถ๐๐ฎ๐น๐ถ๐ด๐ป๐ฒ๐ฑ ๐๐ถ๐๐ต ๐ฒ๐ป๐๐ฒ๐ฟ๐ฝ๐ฟ๐ถ๐๐ฒ ๐ฝ๐ผ๐น๐ถ๐ฐ๐.
๐ง๐ต๐ถ๐ ๐ถ๐ ๐๐ต๐ฒ๐ฟ๐ฒ ๐ฅ๐๐ (๐ฅ๐ฒ๐๐ฟ๐ถ๐ฒ๐๐ฎ๐น ๐๐๐ด๐บ๐ฒ๐ป๐๐ฒ๐ฑ ๐๐ฒ๐ป๐ฒ๐ฟ๐ฎ๐๐ถ๐ผ๐ป) ๐ฐ๐ผ๐บ๐ฒ๐ ๐ถ๐ป
RAG connects LLMs to ๐ฒ๐ป๐๐ฒ๐ฟ๐ฝ๐ฟ๐ถ๐๐ฒ ๐ธ๐ป๐ผ๐๐น๐ฒ๐ฑ๐ด๐ฒ ๐๐ต๐ฒ๐ ๐๐ฒ๐ฟ๐ฒ ๐ป๐ฒ๐๐ฒ๐ฟ ๐๐ฟ๐ฎ๐ถ๐ป๐ฒ๐ฑ ๐ผ๐ป.
Instead of relying only on pre-trained knowledge, it:
1. Breaks documents into ๐ฐ๐ต๐๐ป๐ธ๐
2. Converts them into ๐ฒ๐บ๐ฏ๐ฒ๐ฑ๐ฑ๐ถ๐ป๐ด๐ (numerical representations of text)
3. Stores them in a ๐๐ฒ๐ฐ๐๐ผ๐ฟ ๐ฑ๐ฎ๐๐ฎ๐ฏ๐ฎ๐๐ฒ
4. Uses ๐ฐ๐ผ๐๐ถ๐ป๐ฒ ๐๐ถ๐บ๐ถ๐น๐ฎ๐ฟ๐ถ๐๐(=0.97) to retrieve the most relevant context
5. Feeds that context to the LLM to generate grounded responses
๐๐ต๐๐ป๐ธ๐ถ๐ป๐ด ๐ถ๐ ๐บ๐ผ๐ฟ๐ฒ ๐ถ๐บ๐ฝ๐ผ๐ฟ๐๐ฎ๐ป๐ ๐๐ต๐ฎ๐ป ๐ถ๐ ๐น๐ผ๐ผ๐ธ๐
The way you split data directly impacts the quality of retrieval.
Common chunking strategies include:
โข Fixed-length chunking
โข Sentence-based chunking
โข Paragraph-based chunking
โข Sliding window chunking
โข Semantic chunking
โข Recursive chunking
Better chunking โ Better retrieval โ Better answers
And more importantly:
๐๐ฒ๐๐๐ฒ๐ฟ ๐ฎ๐ป๐๐๐ฒ๐ฟ๐ ๐๐ถ๐๐ต ๐ฐ๐ถ๐๐ฎ๐๐ถ๐ผ๐ป๐
Which means the response is not just helpful, but ๐๐ฟ๐ฎ๐ฐ๐ฒ๐ฎ๐ฏ๐น๐ฒ:
โข You can see where the answer came from
โข You can refer back to the exact company document
โข The answer has real grounding, not just confidence
Thatโs what gives AI ๐ฐ๐ฟ๐ฒ๐ฑ๐ถ๐ฏ๐ถ๐น๐ถ๐๐ ๐ถ๐ป ๐ฒ๐ป๐๐ฒ๐ฟ๐ฝ๐ฟ๐ถ๐๐ฒ ๐๐๐ฒ.
โ๏ธ๐ฅ๐๐ ๐๐ ๐๐ถ๐ป๐ฒ-๐๐๐ป๐ถ๐ป๐ด
๐๐ถ๐ป๐ฒ-๐๐๐ป๐ถ๐ป๐ด โ Good for behavior, structure, and low latency
๐ฅ๐๐ โ Better for dynamic, real-time, and private enterprise data
For most enterprise use cases, RAG is the most effective starting point for improving LLM responses.
๐ ๐ ๐๐ฎ๐ธ๐ฒ๐ฎ๐๐ฎ๐
AI is not just about generating answers.
Itโs about:
โข Retrieving the right context
โข Grounding responses in real data
โข Making answers ๐๐ฒ๐ฟ๐ถ๐ณ๐ถ๐ฎ๐ฏ๐น๐ฒ, ๐ป๐ผ๐ ๐ท๐๐๐ ๐ฐ๐ผ๐ป๐๐ถ๐ป๐ฐ๐ถ๐ป๐ด
A quick shoutout to Arvind Narayanamurthy & Aishwarya Srinivasan from The Gen Academy for this weekend learning.
I ran 7 RAG architectures side by side on the same questions to see what actually matters.
The setup comes from a deep dive by a friend of mine, Vidhyakshaya Kannan, on The Gen Academy which breaks down these RAG architectures and their tradeoffs.
A few things that stood out:
- Most architectures land in a similar accuracy range
- Agentic RAG performs best, but costs ~4.8ร more per correct answer
- Self-RAG underperforms Naive RAG. More complexity, worse results
It is easy to over-engineer RAG systems. Running them side by side makes that tradeoff very clear. Iโll drop the blog post and code in the comments.
How are you deciding between RAG architectures in practice? Optimizing for accuracy, cost, or something else?
#MachineLearning #ArtificialIntelligence #DellTech #DellProPrecision #NVIDIA
I ran 7 RAG architectures side by side on the same questions to see what actually matters.
The setup comes from a deep dive by a friend of mine, Vidhyakshaya Kannan, on The Gen Academy which breaks down these RAG architectures and their tradeoffs.
A few things that stood out:
- Most architectures land in a similar accuracy range
- Agentic RAG performs best, but costs ~4.8ร more per correct answer
- Self-RAG underperforms Naive RAG. More complexity, worse results
It is easy to over-engineer RAG systems. Running them side by side makes that tradeoff very clear. Iโll drop the blog post and code in the comments.
How are you deciding between RAG architectures in practice? Optimizing for accuracy, cost, or something else?
#MachineLearning #ArtificialIntelligence #DellTech #DellProPrecision #NVIDIA
Most AI PMs are managing probabilistic systems with deterministic tools.
That's why AI products quietly rot after launch.
In traditional software: a button either works or it doesn't. Bugs are reproducible. If it passes staging, it passes prod.
In AI: the "same" input produces different outputs. A prompt that works on 10 examples fails on the 11th. Features don't ship as "works / doesn't work", they ship with a distribution of quality.
And most PMs have no idea how to measure that distribution.
I mapped out the full AI PM tooling stack for 2026, the 6 layers every AI PM needs to know, the category leaders in each, and the judgment calls that actually matter:
โ Evaluation (where quality is defined)
โ Observability (seeing what actually shipped)
โ Prompt management (governing changes safely)
โ Prototyping (Bolt, Lovable, v0, Cursor)
โ Experimentation (shipping to users safely)
โ Customer feedback (keeping evals honest)
The biggest mistake AI PMs make? Adopting an eval tool and never building a real dataset. Treating the tool as the deliverable instead of the scaffolding.
Inspired by the Lightning Session on The AI PM Playbook by The Gen Academy, hosted by Aishwarya Srinivasan and Arvind Narayanamurthy, with guest speaker Shyamala Prayaga.
Full breakdown on Substack, link in comments ๐
Most AI PMs are managing probabilistic systems with deterministic tools.
That's why AI products quietly rot after launch.
In traditional software: a button either works or it doesn't. Bugs are reproducible. If it passes staging, it passes prod.
In AI: the "same" input produces different outputs. A prompt that works on 10 examples fails on the 11th. Features don't ship as "works / doesn't work", they ship with a distribution of quality.
And most PMs have no idea how to measure that distribution.
I mapped out the full AI PM tooling stack for 2026, the 6 layers every AI PM needs to know, the category leaders in each, and the judgment calls that actually matter:
โ Evaluation (where quality is defined)
โ Observability (seeing what actually shipped)
โ Prompt management (governing changes safely)
โ Prototyping (Bolt, Lovable, v0, Cursor)
โ Experimentation (shipping to users safely)
โ Customer feedback (keeping evals honest)
The biggest mistake AI PMs make? Adopting an eval tool and never building a real dataset. Treating the tool as the deliverable instead of the scaffolding.
Inspired by the Lightning Session on The AI PM Playbook by The Gen Academy, hosted by Aishwarya Srinivasan and Arvind Narayanamurthy, with guest speaker Shyamala Prayaga.
Full breakdown on Substack, link in comments ๐
Yesterday, I attended the AI PM Playbook by The Gen Academy, featuring Shyamala Prayaga (Senior AI Product Manager at NVIDIA).
You might wonder, whatโs a software engineer doing in an AI PM session?
Because Iโve seen how often we jump straight into buildingโฆ and only later realize we shouldโve asked better questions first. Working closely with PMs, Iโve started valuing that thought process more, how ideas are shaped before they become features. Thatโs what pulled me into this session.
One line from the session that really stayed with me:
๐ Donโt fall for the โshiny object syndromeโ in AI.
Itโs tempting to add AI everywhere, but the real focus should be on:
โข Solving the right problem (not just using AI for the sake of it)
โข Defining metrics & evals upfront - โbuild evals before building featuresโ
โข Thinking about adoption - token economics, context, and real user value
A good reminder that building AI products, at any level, needs clarity, intent, and responsibility, not just hype.
#AI #ProductManagement #SoftwareEngineering #GenAI #Learning
Yesterday, I attended the AI PM Playbook by The Gen Academy, featuring Shyamala Prayaga (Senior AI Product Manager at NVIDIA).
You might wonder, whatโs a software engineer doing in an AI PM session?
Because Iโve seen how often we jump straight into buildingโฆ and only later realize we shouldโve asked better questions first. Working closely with PMs, Iโve started valuing that thought process more, how ideas are shaped before they become features. Thatโs what pulled me into this session.
One line from the session that really stayed with me:
๐ Donโt fall for the โshiny object syndromeโ in AI.
Itโs tempting to add AI everywhere, but the real focus should be on:
โข Solving the right problem (not just using AI for the sake of it)
โข Defining metrics & evals upfront - โbuild evals before building featuresโ
โข Thinking about adoption - token economics, context, and real user value
A good reminder that building AI products, at any level, needs clarity, intent, and responsibility, not just hype.
#AI #ProductManagement #SoftwareEngineering #GenAI #Learning
AI products donโt fail because of bad models.
They fail because of bad decisions.
Over the weekend, I attended the Lightning Lesson on The AI PM Playbook, hosted by Aishwarya Srinivasan and Arvind Narayanamurthy (Maven ร The Gen Academy), with Shyamala Prayaga (NVIDIA).
It exposed a gap in how I was thinking about AI systems. I was focused on capabilities. This shifted me toward accountability.
The MAP Framework (Model โ Augment โ Program) stood outโnot just as a technique, but as a decision filter:
1. When should AI generate?
2. When should it assist?
3. And when should it stay out entirely?
This matters because AI is non-deterministic.
Treat it like traditional software, and you end up with "silent failures"โsystems that appear correct until they break user trust.
The shift for me:
As someone with a technical background, I used to ask: โWhat can this model do?โ Now, as an aspiring AI PM, I ask: โWhat should the system be responsible for?โ
Iโm now applying this product-first mindset to my own projectsโmoving beyond demos toward reliable systems at scale.
For AI builders: whatโs one decision that becomes harder when moving from demo to production?
#ArtificialIntelligence #NVIDIA #AIProductManagement #GenerativeAI #SystemDesign #BuildInPublic #AIEngineering #MachineLearning #ProductStrategy #TheGenAcademy
AI products donโt fail because of bad models.
They fail because of bad decisions.
Over the weekend, I attended the Lightning Lesson on The AI PM Playbook, hosted by Aishwarya Srinivasan and Arvind Narayanamurthy (Maven ร The Gen Academy), with Shyamala Prayaga (NVIDIA).
It exposed a gap in how I was thinking about AI systems. I was focused on capabilities. This shifted me toward accountability.
The MAP Framework (Model โ Augment โ Program) stood outโnot just as a technique, but as a decision filter:
1. When should AI generate?
2. When should it assist?
3. And when should it stay out entirely?
This matters because AI is non-deterministic.
Treat it like traditional software, and you end up with "silent failures"โsystems that appear correct until they break user trust.
The shift for me:
As someone with a technical background, I used to ask: โWhat can this model do?โ Now, as an aspiring AI PM, I ask: โWhat should the system be responsible for?โ
Iโm now applying this product-first mindset to my own projectsโmoving beyond demos toward reliable systems at scale.
For AI builders: whatโs one decision that becomes harder when moving from demo to production?
#ArtificialIntelligence #NVIDIA #AIProductManagement #GenerativeAI #SystemDesign #BuildInPublic #AIEngineering #MachineLearning #ProductStrategy #TheGenAcademy
