New batches starting this week Β· Limited seats

Embeddings Explained: How AI Represents Meaning as Numbers

Embeddings turn text into vectors so that similar meanings sit close together, powering semantic search and RAG. This guide explains how embedding models work, how similarity is measured, how to choose a model and the pitfalls to avoid in production.

Text passes through an embedding model into a vector of numbers used for similarity search and meaning-based results
Last updated Β· 13 min read Β· 2,895 words

An embedding is a list of numbers that represents the meaning of a piece of text, an image or another item, produced by a model trained for that purpose. In a good embedding space, items with similar meanings sit close together, so "how do I return a damaged product?" lands near "refund policy for faulty items" even though the two share almost no words. That single property powers semantic search, retrieval-augmented generation (RAG), clustering, deduplication and recommendations.

What are embeddings? An intuitive definition

Computers cannot compare meanings directly; they compare numbers. An embedding model reads an input, such as a sentence, a paragraph or a product description, and outputs a fixed-length vector: an ordered list of floating-point numbers. Every input to the same model produces a vector of the same length.

Think of an embedding space as a map of meaning with hundreds or thousands of dimensions instead of two. Texts about leave policies cluster in one region, texts about network outages in another, and texts about refunds somewhere else. A question is placed on the same map, and the nearest points are the most relevant passages.

Two clarifications:

  • Individual dimensions usually have no human-readable meaning. There is no "dimension 42 = finance". Meaning is spread across all the numbers, and only the relative positions of vectors matter.
  • Vectors are only comparable within one model. A vector from model A and a vector from model B live on different maps. Comparing them is meaningless, even if they happen to have the same length.

For the tokens and transformers underneath, see what an LLM is.

How embedding models produce vectors

Modern text embedding models are usually transformer encoders: networks that read the whole input at once and build a contextual representation of every token. To get one vector for the whole text, the model pools those token representations, for example by averaging them or by taking the representation of a special token, and often normalises the result to unit length.

What makes the vectors useful is how the model is trained. The common recipe is contrastive learning:

  1. Collect pairs of texts that should be related: a question and the passage that answers it.
  2. Embed both sides and push the vectors of related pairs closer together.
  3. At the same time, push the vectors of unrelated texts apart. Hard negatives, texts that look similar but are not the right answer, teach finer distinctions.

After training on very large numbers of such pairs, the model has learned a geometry where closeness tracks relatedness. Embeddings therefore reflect their training data: a general model may place two legal clauses together because they look alike, even when they mean opposite things.

Some models are asymmetric: they expect queries and documents to be embedded slightly differently, often by adding an instruction or prefix such as "query:" or "passage:" before the text. Skipping that step, when the model's documentation asks for it, quietly lowers retrieval quality.

Similarity measures: cosine similarity and dot product

Once texts are vectors, "how similar are these?" becomes arithmetic. Two measures dominate:

  • Cosine similarity measures the angle between two vectors, ignoring their length. It ranges from -1 to 1; higher means more similar. It is the default for most text embedding models.
  • Dot product multiplies matching components and adds them up. It is cheaper to compute and is affected by vector length. When vectors are normalised to unit length, dot product and cosine similarity give the same ranking and the same values.

The practical rule: use the measure the model's documentation recommends, and configure your vector index with that same measure.

A tiny worked example (toy vectors, purely illustrative)

Real embeddings have hundreds or thousands of dimensions and the values below are invented for teaching. Imagine a 3-dimensional model and three texts:

TextToy vector
A: "refund policy"[0.9, 0.1, 0.2]
B: "how to return an item"[0.8, 0.2, 0.3]
C: "server outage last night"[0.1, 0.9, 0.4]

Dot product of A and B: (0.9 Γ— 0.8) + (0.1 Γ— 0.2) + (0.2 Γ— 0.3) = 0.72 + 0.02 + 0.06 = 0.80.

Length of A: √(0.81 + 0.01 + 0.04) = √0.86 β‰ˆ 0.927. Length of B: √(0.64 + 0.04 + 0.09) = √0.77 β‰ˆ 0.877.

Cosine similarity of A and B: 0.80 Γ· (0.927 Γ— 0.877) β‰ˆ 0.98, very similar.

Dot product of A and C: 0.09 + 0.09 + 0.08 = 0.26. Length of C: √0.98 β‰ˆ 0.990. Cosine similarity of A and C: 0.26 Γ· (0.927 Γ— 0.990) β‰ˆ 0.28, not very similar.

So a search for "refund policy" would rank B above C. Note that the absolute scores mean little on their own; different models produce different score ranges, so a fixed threshold such as "anything above 0.8 is relevant" must be calibrated per model and per dataset, never copied from a tutorial.

Comparing against every stored vector works for thousands of items. At millions, systems use approximate nearest-neighbour indexes; how those work, and how to store vectors with metadata, is covered in Vector Databases Explained.

Dimensions: the size, quality and cost trade-off

The dimension of an embedding is the length of its vector. Models differ widely here, and more dimensions are not automatically better.

FactorLarger vectorsSmaller vectors
Capacity to capture nuanceUsually higherUsually lower, but often enough
Storage per itemMoreLess
Index memory and search latencyHigherLower
Embedding model cost and speedLarger models tend to be slower and costlierSmaller models tend to be faster and cheaper

Two techniques soften the trade-off. Some models are trained so that the first part of the vector carries the most information, letting you truncate vectors to a shorter length with a modest quality loss (sometimes called Matryoshka embeddings). And quantisation stores each number with fewer bits, cutting memory substantially at some cost to precision.

How many dimensions do you need? Measure: compare two or three configurations on your own queries for quality, latency and storage.

Choosing an embedding model

Public leaderboards measure general benchmarks, not your documents. Choose on these criteria, then test.

Domain fit

A model trained largely on general web text may struggle with clinical notes, legal contracts, source code or insurance policy wording. If your corpus is specialised, include domain-specific queries in your evaluation set, and consider fine-tuning an open-weight embedding model on your own query-passage pairs if off-the-shelf quality falls short.

Languages, including Indian languages

Many Indian enterprise use cases involve more than English: customer queries in Hindi, Telugu, Tamil or Kannada, documents in English, and plenty of code-mixed text such as Hinglish or Telugu written in Latin script. Multilingual embedding models are trained to place equivalent meanings in different languages near each other, so a Telugu question can retrieve an English policy passage. Coverage varies by model, and code-mixed text is often weaker. Build test queries that look like what your users actually type, in the scripts they actually use, before committing.

Context length

Every embedding model has a maximum input length in tokens. Text beyond that limit is typically truncated silently, so the end of a long chunk never influences its vector. Match your chunk size to the model's limit; a longer limit does not mean you should embed whole documents, since one vector for a fifty-page file blurs many topics together. See RAG chunking strategies.

Hosted API vs open-weight

  • Hosted embedding APIs (offered through platforms such as Amazon Bedrock, Azure OpenAI and Google Cloud's Vertex AI) need no infrastructure, scale easily and are billed per token. Your text is sent to the provider, so data handling terms and region matter.
  • Open-weight models run on your own servers or cloud account. Data stays inside your network, you control versions, and cost becomes compute rather than per-call fees. You take on serving and upgrades. See Self-Hosting LLMs for the serving side.

Licensing and stability

Open-weight does not always mean free for commercial use; read the licence. For hosted models, check the provider's deprecation policy: if a model version is retired, you will have to re-embed everything (see the pitfalls below).

Choosing and testing embedding models, building vector search and wiring it into a RAG pipeline are core labs in Cloudsoft's AI, GenAI and Agentic AI course.

What embeddings are used for

  • Semantic search. Find documents by meaning rather than exact words. "Work from home rules" finds the "remote working policy".
  • RAG. Embeddings are the retrieval step in most RAG systems: chunks are embedded at indexing time, the user's question is embedded at query time, and the nearest chunks are passed to the LLM as context. The full pipeline is explained in What Is RAG?
  • Clustering. Group similar items without predefined labels, such as recurring themes in support tickets.
  • Deduplication. Find near-duplicate documents, tickets or product listings.
  • Classification. Use embeddings as input features to a lightweight classifier, or compare a new item with labelled examples to route it (for instance, sending a ticket to the right team).
  • Recommendations. Suggest articles, products or courses similar to what a user has viewed, by finding nearby items in the embedding space.

Consider a retailer running customer support in English, Hindi and Telugu. It might cluster past tickets, deduplicate repeats, route new ones and run a RAG assistant over its returns policies. One multilingual embedding model, used consistently, can serve all four tasks; that consistency is exactly what the pitfalls below are about.

Embedding pitfalls in production

Mismatched models between indexing and querying

The most damaging bug is also the quietest. If documents were embedded with one model and queries with another, or with the same model but a different version, prefix or normalisation setting, search still returns results. They are just poor. Nothing errors. Store the model name, version and settings alongside every index, and make your query path read them from that metadata rather than from a separate config value that can drift.

Re-embedding cost when you switch models

Because vectors from different models are incompatible, switching embedding models means re-embedding the entire corpus and rebuilding the index. That needs a migration plan: build the new index alongside the old one, evaluate both on the same queries, then switch traffic. Keep source text and chunk boundaries so re-embedding skips re-ingestion.

Chunk quality

An embedding can only represent what is in its chunk. A chunk that cuts a rule in half or flattens a table into noise embeds poorly with any model. Adding the document title and section heading to each chunk before embedding often helps the vector capture context.

Numbers, codes and exact terms

Embeddings capture general meaning well and exact strings poorly. A query for policy number "HL-2291", error code "ORA-01017" or a specific product SKU may retrieve passages about similar-looking codes, because to the model they mean roughly the same thing. Negation and numeric ranges can blur too. This is why production systems usually combine vector search with keyword search (such as BM25) in hybrid search, and then rerank the merged results with a cross-encoder that reads the query and passage together. Our guide to hybrid search and reranking in RAG explains how to set this up.

Evaluating retrieval quality

You cannot judge an embedding model by looking at vectors. You judge it by whether the right passages come back for real questions. A practical approach:

  1. Build a test set of representative queries, each labelled with the chunk or chunks that should be retrieved. Include paraphrases, typos, exact-code lookups and, where relevant, Indian-language and code-mixed queries.
  2. Measure retrieval metrics such as recall at k (did the right chunk appear in the top k?) and mean reciprocal rank (how high did it appear?).
  3. Compare candidate models, dimensions and chunking settings on the same set, alongside latency and cost.
  4. Re-run it whenever the model, chunking or index settings change.

Retrieval metrics are only half the picture for RAG; you also need to check whether the final answers are faithful to the retrieved context. Our guide to RAG evaluation metrics covers both retrieval and generation metrics and tools such as Ragas.

Privacy considerations

It is tempting to treat embeddings as anonymised data because they look like meaningless numbers. They are not. Research has shown that a surprising amount of the original text can be reconstructed from embeddings, so treat a vector index with the same sensitivity as the documents it was built from.

  • Where the text goes. With a hosted embedding API, every chunk and every user query is sent to the provider. Check retention, training use and processing region.
  • Access control. Store permissions as metadata with each vector and filter at query time, so a user can only retrieve chunks they are allowed to see.
  • Deletion. When a person's data must be removed, for example under India's Digital Personal Data Protection Act, the matching vectors must be removed too, not just the source files.
  • Minimise before embedding. Mask or remove personal data that retrieval does not need, such as account numbers or Aadhaar numbers, before chunks are embedded.

A hospital indexing discharge summaries, for example, might self-host an open-weight model, filter by department and redact patient identifiers first. For wider controls, see AI Security for Enterprises.

Making these choices inside a customer's environment, with their data, permissions and compliance rules, is the kind of work Forward Deployed Engineers do; Cloudsoft's FDE PRO program builds it into its Enterprise Knowledge Assistant project using PostgreSQL with pgvector.

Frequently asked questions

What are embeddings in simple terms?

Embeddings are lists of numbers that represent the meaning of text, images or other items. A model converts each item into a vector so that items with similar meanings end up close together.

Is an embedding model the same as an LLM?

No. An embedding model turns input into a vector for comparison and search; it does not generate text. An LLM generates text. A RAG system usually uses both: an embedding model to retrieve relevant passages and an LLM to write the answer from them.

Do more dimensions mean better embeddings?

Not necessarily. Larger vectors can capture more nuance but cost more storage, memory and search time. A smaller model often does as well on a given dataset, so test on your own queries.

Can I mix embeddings from two different models in one index?

No. Vectors from different models live in different spaces and cannot be meaningfully compared, even if they have the same length. If you switch models you must re-embed the whole corpus and rebuild the index.

Do embeddings work for Hindi, Telugu and other Indian languages?

Multilingual embedding models can represent Indian languages and can match a question in one language with a passage in another. Quality varies by model and language, and code-mixed or transliterated text is often harder. Evaluate with queries written the way your users actually write them.

Why does semantic search miss exact codes and numbers?

Embeddings capture general meaning, so similar-looking codes or numbers can land close together even when they refer to different things. Combining vector search with keyword search in a hybrid setup, then reranking the results, fixes most of these misses.

Getting embeddings right is mostly careful engineering: consistent models, good chunks, hybrid retrieval and honest evaluation. If you want to build these systems end to end with guided labs, explore Cloudsoft's AI, GenAI and Agentic AI training in Hyderabad, available in our Ameerpet classroom beside Ameerpet Metro or live online. Call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us