House of Data/Curriculum
What you actually learn
The whole syllabus, in the open.
No "advanced modules" you only hear about after you pay. Fifteen modules in the order they become useful — Python and SQL, transformers and embeddings, RAG and document extraction, agents and orchestration, Claude Code, and the cloud services that run it all — eight levels from your first line of code to your first offer, and the full tool list.
The syllabus
Fifteen modules, in the order they become useful.
Foundations, then the models, then building, then production, then the job. No module starts before the one under it is solid, and every bullet below is taught — there is no separate "advanced" course.
Where it all sits
Generative AI, machine learning, deep learning and data science — the map
- What machine learning, deep learning and Generative AI actually are, and where data science sits among them
- Supervised and unsupervised learning in one sitting; why neural networks changed the ground
- What a model can and cannot do, and the vocabulary an interview panel expects you to have
Python, and vibe coding
The Python you use daily, and building alongside an AI rather than instead of one
- Python that earns its place: data structures, functions, files, HTTP calls, async basics, pandas
- Environments, packages and notebooks against scripts — how a project is actually laid out
- Vibe coding: building with an AI pair, reading what it wrote, and knowing when to stop trusting it
- Excel where Excel is still the faster answer
SDLC and AIDLC
How software gets built, and how an AI product differs
- Requirements, design, build, test, release, run — and who owns each part
- The AI development life cycle: datasets, evaluation sets, prompt versions, model swaps, drift
- Agile delivery, tickets, code review and what “done” means on a team
Data — structured and not
DBMS, how SQL really works, and everything that is not a table
- DBMS concepts: tables, keys, relationships, normalisation, transactions
- SQL properly — joins, grouping, window functions, and what the query planner does with your query
- NoSQL where it fits
- Unstructured data: text, PDFs, scans, images and audio — parsing, cleaning, chunking, and handling personal data
Inside the models
Transformer architecture, LLMs, SLMs and embedding models
- Transformer architecture: attention, tokens, context windows, why length costs money
- LLMs against SLMs, hosted against open-weight, quantisation and where each one belongs
- Embedding models, similarity, reranking — the machinery under every retrieval system
- Fine-tuning, LoRA and PEFT, and the more useful skill of knowing when not to fine-tune
- Multimodal models: vision, speech and documents
Prompting, context and evaluation
Getting reliable output, and proving that it is reliable
- Prompt patterns, system prompts, structured output and function schemas
- Context engineering, caching, and keeping token cost under control
- Evaluation sets, LLM-as-judge, regression testing a prompt, catching hallucination
- Guardrails, red teaming, responsible AI, PII handling and AI security
RAG, and documents
Vector databases, retrieval, document extraction, chatbots and KAG
- Chunking strategies, embeddings, hybrid search, reranking, citations
- Vector databases: Pinecone, Qdrant, FAISS, ChromaDB, pgvector and Azure AI Search
- Document extraction and document intelligence — OCR, tables, forms, Amazon Textract and Azure Document Intelligence
- The types of chatbot: scripted, RAG, tool-using, multi-agent and voice — and which one a business actually needs
- KAG and knowledge graphs, GraphRAG, and where a graph beats a vector
- Measuring retrieval instead of hoping: recall, groundedness, answer quality
Agents and orchestration
LangChain, LangGraph, LangSmith, MCP, middleware agents and n8n
- What an agent really is: tools, memory, planning, loops and stopping conditions
- LangChain and LangGraph for orchestration; LangSmith for tracing and evaluation
- MCP, tool calling and middleware agents — how an agent reaches your systems safely
- Multi-agent systems and agentic AI patterns, with a human in the loop where it matters
- n8n and workflow automation for the work that does not need a model at all
- Cost, latency, retries and what to do when an agent goes round in circles
Coding with AI tools
GitHub Copilot, Cursor and Amazon Kiro, used the way teams use them
- GitHub Copilot, Cursor and Amazon Kiro — what each is good at
- Spec-driven development, giving a tool the right context, and reviewing what comes back
- Where these tools break, and the habits that keep a codebase yours
Claude, in depth
Claude Code and Cowork — a separate module, because this is where the work is going
- Claude Code: agentic coding in the terminal, subagents, hooks, skills and plugins
- Cowork: automating file and task work, and building your own skills for it
- The Claude API: tool use, MCP servers, prompt caching and batch
The product around the model
APIs, GitHub, front end and UX — the parts that make it a product
- Git and GitHub: branching, pull requests, reviews and Actions
- REST, FastAPI, OpenAPI and Swagger, tested with Postman
- Authentication, rate limiting, streaming responses and webhooks
- HTML, CSS and JavaScript basics, then React or Angular fundamentals
- UX design and Figma — wireframe to handoff
- Streamlit and Gradio when a demo is all you need
Containers and pipelines
Docker, Kubernetes, Jenkins and CI/CD
- Docker: images, containers, volumes and Compose
- Kubernetes: pods, deployments, services, scaling and health checks
- Jenkins and GitHub Actions: build, test, deploy, roll back
- Environments, secrets, logging and monitoring — MLOps and LLMOps in practice
Cloud, and its AI services
AWS, Azure, GCP and Cloudflare — the services that actually appear in job descriptions
- AWS: S3, ECS, EKS, SQS, Lambda, Bedrock, Textract, SageMaker and OpenSearch
- Azure: Blob Storage, Key Vault, Document Intelligence, AI Search, Azure OpenAI, Functions and Azure DevOps
- GCP: Vertex AI, Document AI, Cloud Run and BigQuery
- Cloudflare: Workers, Pages, R2, D1, Vectorize and Workers AI
- IAM, regions, quotas and what each of these costs to leave running
The projects
Two mainstream real-time projects, an end-to-end RAG system and a document extraction pipeline
- Two mainstream real-time projects, scoped the way a team would scope them
- An end-to-end RAG system: ingestion, retrieval, evaluation, interface, deployment
- A document extraction project — messy PDFs and scans in, structured data out
- All of it versioned, deployed and monitored, with a public link you can hand to an interviewer
Get placed
Profile, applications and interviews — until you are working
- A resume written around these projects, not a template
- LinkedIn and Naukri rebuilt, direct submissions to hiring partners
- Mock interviews repeated until the answers hold up
The build sequence
Eight levels. Foundation to roof.
Nothing here is optional and nothing is left to you. Each level is scheduled, taught and checked before the next one starts.
Enrol
Register, pick your batch and get a profile analysis — graduation year, gaps, prior experience — with a realistic package report before you spend a rupee on the course.
Live training starts
Online, offline, weekday or weekend. Fifteen students maximum, so you can interrupt and ask.
Real-time projects
Two mainstream real-time projects, an end-to-end RAG system and a document extraction pipeline — built the way they're built at work, versioned, deployed to Azure, GCP, AWS or Cloudflare, with MLOps and LLMOps around them.
One-to-one support
Stuck at 11pm on a model that won't converge? That's what the mentor line is for. Any time, through the program.
Corporate readiness
Interview vocabulary, corporate etiquette, team outings. The part most institutes skip and every interview panel notices.
Certification
Complete the program and get certified by House of Data, with your project portfolio attached.
Profile marketing
We build the resume ourselves and push your profile through Naukri, LinkedIn and our hiring partners, tuned to whatever the market is asking for that month.
Placed
Multiple offers, your choice of city and company type — product, service, hybrid or startup. If it doesn't happen here, you move to the Advanced Placement Plan.
What you will have built
A retrieval pipeline, end to end.
Not a notebook that calls an API. A system that finds the right passage, hands it to the model, and answers from your documents rather than from memory — deployed, measured and running on a public URL.
This is the diagram on the whiteboard in the photo above. By the end of the program it is also the diagram of something you have shipped.
The stack
The tools on the job descriptions.
Not a survey of everything that exists. This is what Indian hiring managers are asking for in 2026, and it's what you'll have your hands on during the program. Gold means it's showing up in almost every Gen AI job description right now.
Languages & data
- Python
- SQL
- Pandas
- NumPy
- NoSQL
- Advanced Excel
- Unstructured data
- Chunking
Models & Gen AI
- Transformers
- LLMs
- SLMs
- Embedding models
- Hugging Face
- OpenAI API
- Anthropic API
- Gemini API
- Fine-tuning & LoRA
- Multimodal
RAG & vector data
- RAG
- Pinecone
- Qdrant
- FAISS
- ChromaDB
- pgvector
- Azure AI Search
- Reranking
- KAG
- Knowledge graphs
- Amazon Textract
- Azure Document Intelligence
Agents & orchestration
- LangChain
- LangGraph
- LangSmith
- MCP
- Tool calling
- Multi-agent systems
- CrewAI
- n8n
- Agentic AI
Building with AI
- Claude Code
- Cowork
- GitHub Copilot
- Cursor
- Amazon Kiro
- Vibe coding
- Spec-driven development
APIs & the front end
- FastAPI
- REST
- Swagger / OpenAPI
- Postman
- Git & GitHub
- HTML
- CSS
- JavaScript
- React
- Angular
- Figma
- Streamlit
Deployment
- Docker
- Kubernetes
- Pods & images
- Jenkins
- GitHub Actions
- CI/CD
- MLOps
- LLMOps
Cloud services
- AWS
- EKS
- ECS
- SQS
- S3
- Bedrock
- Lambda
- Azure
- Blob Storage
- Key Vault
- Azure DevOps
- GCP
- Vertex AI
- Cloudflare
Evaluation & governance
- LLM evaluation
- LLM-as-judge
- Guardrails
- Red teaming
- Responsible AI
- PII handling
- AI security
- Cost & token control
Ways of working
- SDLC
- AIDLC
- Agile delivery
- Code review
- System design
- Interview vocabulary
Two lengths, one syllabus.
Everything above is taught on both programmes — the Fast Track compresses it into 2 + 1 months, the Flagship takes twenty-four weeks over it and goes deeper on agents, LLMOps and your own project.