New batches starting this week · Limited seats

How We Built Aanya AI: Inside Cloudsoft's Self-Hosted RAG Assistant

Behind the scenes of Aanya AI, Cloudsoft's self-hosted RAG assistant: the architecture, structure-aware chunking, hybrid (dense + BM25) retrieval, embeddings and the guardrails that keep every answer grounded and honest.

Edge AI interview questions 2026: 50 questions on quantisation, runtimes, TinyML, small LLMs on devices, OTA updates and model security
Last updated · 4 min read · 980 words

Aanya AI is the friendly assistant now live on cloudsoftsol.com — your AI Career Guide. Ask her about courses, fees, placement support, formats or a free demo, and she answers in seconds, strictly from verified Cloud Soft content, and hands you to a human counsellor when that's the right move. This post is a behind-the-scenes look at how Aanya actually works: the Retrieval-Augmented Generation (RAG) architecture, how we chunk and index our knowledge, how retrieval picks the right facts, and the guardrails that keep her honest. It's also a real-world example of exactly what we teach in our FDE PRO and APEX programs.

What is Aanya, in one line?

Aanya is a self-hosted RAG assistant: a large language model (LLM) that doesn't answer from memory, but from a curated knowledge base of Cloud Soft's own pages and FAQs. Every answer is grounded in retrieved facts and cites its sources — so she can be helpful without making things up.

The architecture at a glance

Aanya runs entirely on our own infrastructure — no customer data is sent to a third-party AI provider. The core pieces:

  • Vector database (Qdrant) — stores the knowledge base as searchable embeddings.
  • Embedding model — turns text into vectors so we can search by meaning, not just keywords.
  • A compact, locally-hosted LLM — generates the final answer from the retrieved context.
  • An API + guard layer — orchestrates retrieval, assembles context, runs safety checks, and streams the reply.
  • A lightweight web widget — the chat bubble you see, isolated so it never interferes with the site.

The request flow is simple: your question → detect intent → retrieve the most relevant facts → build a tight context → the LLM drafts an answer → guardrails verify it → stream the reply with sources.

Building the knowledge base: ingestion & chunking

A RAG system is only as good as how its knowledge is prepared. We ingest Cloud Soft's authoritative pages, course and program details, and a hand-curated FAQ. The key step is structure-aware chunking — instead of dumping whole pages into the index, we split content into small, self-contained pieces:

  • Each FAQ question–answer pair becomes its own chunk, so a specific question retrieves a specific answer.
  • Long pages are split by section, targeting a few hundred tokens per chunk with a little overlap.
  • Every chunk carries a contextual header (page title, section, program) and metadata — tier, category, and the programs/courses it covers.
  • Curated facts override crawled text. Fees, durations, contact details and placement wording come only from an owner-approved source of truth, so the bot can never drift to outdated copy.

Each chunk is embedded and upserted into the vector index with a deterministic ID, so re-indexing updates changed content and removes stale chunks instead of creating duplicates.

Retrieval: hybrid search that actually finds the right fact

When you ask a question, Aanya runs hybrid retrieval — two searches at once:

  • Dense (semantic) search matches by meaning, so "how much does it cost?" finds the fees even without the word "fee".
  • Sparse (BM25 keyword) search nails exact terms like "FDE PRO", "Intune" or "Entra ID".

The two result sets are fused with Reciprocal Rank Fusion (RRF) and filtered by metadata — for example, fee and placement questions are answered only from authoritative sources. We deliberately keep the retrieved context tight (the top few chunks): a focused context gives faster, more accurate answers than stuffing everything in.

Generation + guardrails: helpful, never fabricated

The LLM receives the curated facts and retrieved context as data (not instructions), plus a system prompt that forbids inventing anything. Then a deterministic guard layer checks the draft before you ever see it:

  • Grounding check — every amount, percentage, phone number and date in the answer must appear in the retrieved context, or the answer is regenerated/handed off.
  • Policy wording — placement is always described as support until you're placed, never as a "guarantee".
  • Anti-hallucination guards — she won't invent branches in other cities (we have one centre in Ameerpet, plus live online).
  • Source citations + human handoff — answers link to real pages, and high-intent or uncertain queries offer a counsellor callback.

If confidence is low, Aanya says she doesn't have that detail and connects you to the team — a graceful, honest fallback instead of a confident wrong answer.

Why self-hosted?

Running the model and vector database on our own servers keeps prospective-student conversations private, avoids per-query API costs, and gives us full control over quality. It's also a brilliant teaching stack — the exact tools used in modern enterprise AI.

Want to build assistants like Aanya?

Everything above — embeddings, vector databases, hybrid retrieval, RAG, agents and guardrails — is hands-on curriculum in our 2026 AI programs:

Frequently asked questions

What is RAG (Retrieval-Augmented Generation)?

RAG is an architecture where the model retrieves relevant facts from a knowledge base and generates its answer from them, instead of relying on what it memorised during training. It makes answers accurate, current and citable.

Why does Aanya cite sources and sometimes hand off to a counsellor?

Because she only answers from verified Cloud Soft content. When a question needs a human — batch dates, personalised guidance, or anything she can't ground in the data — she offers a callback instead of guessing.

Can I learn to build a RAG assistant at Cloud Soft?

Yes — the FDE PRO and APEX programs teach embeddings, vector search, RAG and agents with real projects.

Try Aanya yourself — she's the chat bubble on this site. Or talk to our team: call/WhatsApp +91 96660 19191 or book a free demo.

New · AI Career Guide

Meet Aanya — ask anything about courses, fees & placement

Instant answers from verified Cloudsoft info — courses, fees, formats, placement support and free demos. Available 24/7, right here on the site.

How Aanya works →
Share𝕏inf✉
EnrollWhatsAppCall us