Knowledge graph interview questions in 2026 test whether you can model a messy business domain as entities and typed relationships, query it correctly, keep it accurate as sources change, and decide honestly when a graph is worth its cost compared with a relational table or a vector index. This guide collects 50 high-value questions with model answers, from property graphs, RDF and ontologies through Cypher, GQL and SPARQL, entity resolution, graph algorithms, GraphRAG, provenance, access control and scaling, ending with ten production scenarios. It suits data engineers, AI engineers and backend developers preparing for knowledge graph, graph database and GraphRAG interviews.
How to use this guide
The mechanics of GraphRAG (LLM extraction, community summaries, text-to-Cypher, a reference architecture) are explained in our article on GraphRAG and knowledge graphs, so this page goes deeper on the graph itself: modelling, querying, quality and operations. What interviewers commonly probe at each level:
- Freshers: nodes, edges and properties, property graph versus RDF, a simple Cypher
MATCH, and why a join-heavy question suits a graph. - Mid-level engineers: schema design, entity resolution, idempotent loading, variable-length paths, PageRank and community detection, and GraphRAG local versus global search.
- Senior engineers and architects: provenance and confidence models, access control in the query layer, supernodes and partitioning, graph versus vector trade-offs, and how to recover from bad merges in production.
Practise writing the queries by hand. Interviewers often ask you to write a two-hop Cypher pattern on a whiteboard and then ask what happens when one node has a million edges.
- Fundamentals: graphs, RDF and ontologies (Q1βQ9)
- Query languages: Cypher, GQL, SPARQL, Gremlin (Q10βQ15)
- Building knowledge graphs (Q16βQ21)
- Graph algorithms, embeddings and GNNs (Q22βQ27)
- GraphRAG and graphs versus vectors (Q28βQ33)
- Quality, provenance, security and scale (Q34βQ40)
- Real-world scenario questions (Q41βQ50)
- Key takeaways
- Interview preparation checklist
- FAQ
Fundamentals: graphs, RDF and ontologies
1. What is a knowledge graph, and how is it different from just "a graph database"?
Answer: A knowledge graph is a model of real-world entities and the typed relationships between them, with agreed meaning: what a "Supplier" is, what "SUBSIDIARY_OF" implies, which source each fact came from. A graph database is the storage and query engine. You can run a graph database with no shared semantics (a social graph of user IDs), and you can hold a knowledge graph in RDF files, a triple store or even relational tables. What makes it a knowledge graph is the combination of a schema or ontology, identity (one node per real-world thing), and facts that are integrated from several sources and can be traced back to them.
Interview tip: Say "graph database is the engine, knowledge graph is the curated model of meaning". It shows you know the hard part is semantics and identity, not storage.
2. Explain nodes, edges and properties with a concrete example.
Answer: Nodes are the entities, usually with one or more labels such as Customer or Account. Edges are typed, directed relationships such as (Customer)-[:OWNS]->(Account). Properties are key-value attributes on either: name and kycStatus on the customer, since and role on the OWNS edge. The key modelling decision is what becomes a property and what becomes a node. If you ever need to traverse through a value, share it between entities or attach facts to it, it should be a node. A phone number stored as a property cannot easily reveal that ten customers share it; a Phone node connected to all ten makes that one hop.
3. What is the difference between a property graph and RDF?
Answer: A labelled property graph stores nodes and edges, both of which can carry properties, and edges have their own identity. It is application-oriented and queried with Cypher, GQL or Gremlin. RDF (Resource Description Framework, a W3C standard) represents everything as subject-predicate-object triples, where subjects and predicates are IRIs, globally unique identifiers. It is designed for data integration across organisations, has formal semantics through RDFS and OWL, and is queried with SPARQL. The practical differences: property graphs make edge attributes natural, while classic RDF needs reification or named graphs for statements about statements (RDF 1.2 work on triple terms addresses this). RDF makes shared vocabularies, linking across datasets and reasoning natural.
| Aspect | Property graph | RDF |
|---|---|---|
| Unit of data | Nodes, edges, properties | Triples (subject, predicate, object) |
| Identity | Internal IDs plus your keys | Global IRIs |
| Edge attributes | Native | Reification, named graphs or RDF 1.2 triple terms |
| Schema and reasoning | Constraints, optional schema | RDFS, OWL, SHACL validation |
| Query language | Cypher, GQL, Gremlin | SPARQL |
| Typical fit | Fraud, recommendations, GraphRAG, IT dependencies | Life sciences, publishing, regulatory and cross-organisation data |
4. What is the difference between a taxonomy, a schema and an ontology?
Answer: A taxonomy is a hierarchy of categories with "is a kind of" or "narrower than" links, such as product categories or an ICD-style diagnosis tree. A schema describes structure: which node labels, relationship types and properties exist and which constraints apply (unique customer ID, required country). An ontology goes further and defines meaning: classes, relationships with domains and ranges, rules such as "SUBSIDIARY_OF is transitive" or "a Person cannot be a Legal Entity", so that software can check consistency and infer new facts. A taxonomy is often one part of an ontology. In interviews, give an example: "Laptop is narrower than Electronics" is taxonomy; "every Order must have exactly one Customer" is schema; "if A owns more than half of B, A controls B" is ontology logic.
5. What are RDFS, OWL and SHACL used for?
Answer: RDFS gives basic vocabulary for RDF: classes, subclasses, properties, domains and ranges. OWL (Web Ontology Language) adds richer logic: equivalence, disjointness, cardinality, transitive and inverse properties, so a reasoner can infer facts and detect contradictions. OWL works under the open-world assumption: a missing fact is unknown, not false, which is why OWL is poor at "every customer must have an email" checks. SHACL (Shapes Constraint Language) fills that gap. It validates RDF data against shapes, closed-world style, and reports violations. A common production pattern is OWL or a light RDFS ontology for meaning and SHACL for data quality gates in the load pipeline.
6. Why do graphs handle multi-hop questions better than relational joins?
Answer: In a relational database each hop is a join that the engine resolves through indexes, and a recursive or variable-depth question needs recursive CTEs whose cost and readability degrade with depth. Native graph engines store adjacency directly, so following an edge from a node costs roughly the same regardless of total graph size (often called index-free adjacency, though implementations vary). The query also reads like the question: (a)-[:OWNS]->()-[:TRANSFERS_TO*1..4]->(b). Be fair to SQL, though: for fixed one- or two-hop joins with good indexes, PostgreSQL is fast and simpler to run. Graphs win when depth is variable, paths matter, or the schema of relationships keeps evolving.
7. When would you NOT use a knowledge graph?
Answer: When the questions are aggregations over large flat tables (a warehouse does that better), when relationships are few and fixed (a couple of foreign keys), when the answer lives inside one document (vector or hybrid search is cheaper), or when nobody owns entity identity and the schema. A graph without ownership turns into a duplicated, contradictory "hairball" that is harder to trust than the source systems. Also be cautious if the team has no graph skills and the use case is a single dashboard; the operational cost is real.
Interview tip: Interviewers like candidates who can argue against their own favourite technology. Name the cheaper alternative for each case.
8. Name some graph databases and describe how they differ, without ranking them.
Answer: Neo4j is a widely used native property graph database and the origin of Cypher, with a graph data science library for algorithms. Amazon Neptune is a managed AWS graph service that supports property graphs (queried with Gremlin and openCypher) and RDF (queried with SPARQL), though a given graph's data is accessed through one model, not both; Neptune Analytics is a separate in-memory engine aimed at analytics and algorithms. Other options include Memgraph (in-memory, Cypher-compatible), TigerGraph (distributed, its own GSQL language), ArangoDB (multi-model), JanusGraph (open-source, Gremlin, pluggable storage), Azure Cosmos DB for Apache Gremlin, triple stores such as GraphDB and Apache Jena, and Apache AGE, which adds openCypher queries to PostgreSQL. The choice depends on data model, query language, scale, managed versus self-run, and what your team and cloud already support. Check current documentation for features, since this space moves quickly.
9. How do you design a graph schema for a new domain?
Answer: Start from the questions, not the data. Write ten to twenty real questions the business needs answered, then sketch the nodes and relationship types needed to answer each as a path. Rules that work well: use specific relationship types (SUPPLIES_PART, not RELATED_TO); model events as nodes when they have their own attributes or participants (a Transaction node, not just a TRANSFERRED edge); keep stable business keys and unique constraints on every node label; put time on edges or event nodes so you can query "as of" a date; and keep a schema document that maps every label and type to its source system and owner. Then load a sample and test the questions before loading everything.
Real-world example: Consider a hospital modelling referrals. Modelling Referral as a node, linked to the patient, the referring doctor, the receiving department and the diagnosis, lets you ask questions about referral chains that a simple REFERRED_TO edge cannot answer.
Query languages: Cypher, GQL, SPARQL, Gremlin
10. Write a Cypher query to find customers who share a phone number with a customer flagged for fraud.
Answer: Match the pattern through the shared node, exclude the flagged customer itself, and return distinct results with a limit:
MATCH (f:Customer {flagged: true})-[:HAS_PHONE]->(p:Phone)
<-[:HAS_PHONE]-(c:Customer)
WHERE c <> f
RETURN DISTINCT c.id, p.number, f.id AS flaggedId
LIMIT 100
Points interviewers look for: pattern matching with direction, filtering on a property in the node pattern, DISTINCT because a customer may match through several flagged neighbours, and a LIMIT. Mention that c <> f is needed even though Cypher does not reuse the same relationship twice in one pattern, because c and f are different variables that could bind to the same node through different edges if a customer has two HAS_PHONE edges to the same phone. Also mention parameters ($id) instead of string concatenation.
11. What is the difference between MATCH, MERGE and CREATE in Cypher, and what goes wrong with MERGE?
Answer: CREATE always creates. MATCH finds existing patterns. MERGE matches the whole pattern and creates the whole pattern if it does not exist. The classic bug is merging a full path: MERGE (c:Customer {id:$c})-[:OWNS]->(a:Account {id:$a}) creates a duplicate customer and account if the edge does not exist yet, because the whole pattern did not match. The safe pattern is to merge each node on its key first, then merge the relationship:
MERGE (c:Customer {id: $c})
MERGE (a:Account {id: $a})
MERGE (c)-[r:OWNS]->(a)
ON CREATE SET r.since = $since
Back it with a uniqueness constraint on Customer.id and Account.id, otherwise concurrent loaders can still create duplicates.
12. How do you write variable-length path and shortest path queries, and what are the risks?
Answer: A variable-length pattern such as -[:DEPENDS_ON*1..5]-> follows one to five hops. A shortest path between two known nodes can be written as:
MATCH p = shortestPath(
(a:Account {id: $from})-[:TRANSFER*..6]-
(b:Account {id: $to}))
RETURN p
Recent Neo4j versions also support the GQL-style SHORTEST path selectors. The risks: unbounded patterns (* with no upper limit) can explode combinatorially, undirected patterns double the work, and a path through a supernode can touch millions of edges. Always set an upper bound, restrict relationship types, filter early, and use PROFILE to read the plan. For weighted shortest paths (cost, latency, distance), use an algorithm such as Dijkstra from a graph algorithms library rather than hop counts.
13. What is GQL and how does it relate to Cypher?
Answer: GQL (Graph Query Language) is an ISO/IEC international standard for querying property graphs, published in April 2024 as ISO/IEC 39075. It is the first new ISO database query language since SQL, and its pattern-matching syntax draws heavily on Cypher (and on related work such as the SQL/PGQ extension, which adds property graph queries to SQL). A basic GQL query looks familiar to Cypher users:
MATCH (c:Customer)-[:OWNS]->(a:Account)
WHERE c.id = 'C1001'
RETURN a.number
Vendors are adopting GQL at different speeds, and support for specific features varies, so check each database's conformance notes. The value for a team is portability of skills and, over time, of queries.
14. Write a SPARQL query, and explain property paths.
Answer: SPARQL matches triple patterns with variables. To find suppliers whose parent company, at any depth, is on a watch list:
PREFIX ex: <http://example.org/ns#>
SELECT DISTINCT ?supplier ?parent
WHERE {
?supplier a ex:Supplier ;
ex:subsidiaryOf+ ?parent .
?parent ex:onWatchList true .
}
The + is a SPARQL 1.1 property path meaning "one or more hops"; * means zero or more, / chains predicates and ^ inverts one. a is shorthand for rdf:type. Mention that if a reasoner is enabled, the same query can also return results implied by the ontology (for example, subclasses of Supplier), which is a key difference from property graph querying.
15. How does Gremlin differ from Cypher, and why does an interviewer care?
Answer: Cypher and GQL are declarative: you describe the pattern and the planner decides how to find it. Gremlin (Apache TinkerPop) is a traversal language: you write the steps, such as g.V().has('Customer','id','C1001').out('OWNS').values('number'). Gremlin gives fine control and works across many TinkerPop-compatible engines, but performance depends more on how you order the steps, for example filtering selective properties first. An interviewer asks this to check whether you understand query planning: in a declarative language you read the plan and add indexes; in an imperative traversal you are effectively writing the plan.
Building knowledge graphs
16. Walk me through a pipeline for building a knowledge graph from enterprise data.
Answer: Load structured sources first, because they give the graph its trusted backbone: ERP vendors, CRM customers, CMDB services, HR org charts, each with stable keys. Then enrich from unstructured text (contracts, tickets, reports) through extraction. Every step writes provenance.
structured (ERP, CRM, CMDB) --> map to schema --+
|
docs --> parse --> extract (NER, LLM) ----------+
v
entity resolution + SHACL/rules
v
graph DB (with provenance)
v
quality checks, review queue, metrics
Design for incremental updates from day one: change data capture or event feeds for structured sources, document hashes for text, and idempotent upserts so a re-run does not duplicate anything.
17. What is entity resolution, and how do you do it at scale?
Answer: Entity resolution decides which records refer to the same real-world thing: "Ramesh K. Reddy, Kukatpally" in the CRM and "R. K. Reddy" in a loan file. The standard pipeline is: normalise (case, punctuation, legal suffixes such as Pvt Ltd, transliteration variants); block, so you only compare plausible pairs (same PIN code, same phonetic key, same embedding cluster); score pairs with rules, string similarity, embeddings or a trained classifier; then cluster matched pairs into entities, watching for transitive chains (A matches B, B matches C, but A and C are clearly different). Strong identifiers such as PAN, GSTIN or a vendor code should dominate weak signals such as name similarity. Keep both thresholds: auto-merge above a high score, send the grey zone to human review, and never merge below a floor.
Interview tip: Mention that you never destroy source records. You link them to a resolved entity, so a wrong merge can be undone (see Q41).
18. How do you do relation extraction from text?
Answer: Relation extraction finds typed relationships between entities in text, such as "Acme Components supplies brake pads to Plant 3". Options range from rules and patterns (precise, brittle), to fine-tuned classifiers that label entity pairs in a sentence (good for high-volume, stable relation types), to LLM extraction with a schema (flexible, quick to start, needs validation). In all cases you first detect and link the entities, then decide the relation type, its direction and its qualifiers, such as dates or quantities. Evaluate per relation type with precision and recall on a labelled sample, because overall numbers hide weak types. For regulated domains, keep the source sentence on the edge so a reviewer can verify it.
19. How would you use an LLM for knowledge graph extraction safely?
Answer: Give the LLM a closed schema (allowed entity types, relationship types, properties and their directions) and request structured output that is validated against a JSON schema, as described in our guide to function calling and structured outputs. Then validate before anything reaches the graph:
- Reject unknown types instead of letting the model invent
RELATED_TOorPARTNERS_WITHvariants. - Require an evidence span from the source chunk and check that it actually appears in the text.
- Check domain and range: a
SUBSIDIARY_OFedge must connect two organisations. - Route new entities through entity resolution, never straight into the graph.
- Mark LLM-derived edges with extraction run, model and prompt version, and a confidence or review status.
Measure precision and recall against a hand-labelled set of documents, and rerun that set whenever the prompt or model changes.
20. How do you handle time in a knowledge graph?
Answer: Most enterprise facts are true for a period: a director sits on a board from one date to another, a service depends on a database until a migration. Options are validity properties on edges (validFrom, validTo), event nodes (an Appointment node with dates), or snapshots of the graph. Distinguish valid time (when it was true in the world) from transaction time (when the graph learned it), especially for audits: "what did we know when we approved this loan?" Queries then filter with an "as of" parameter. Never delete a relationship that ended; close it, or you lose history that compliance and incident reviews need.
21. How do you keep a knowledge graph in sync with its source systems?
Answer: Treat the graph as a derived store with a clear source of truth per fact. For structured systems, use change data capture or event streams and idempotent upserts keyed on business IDs; for deletes, close or remove the edges whose provenance points only to the deleted record. For documents, hash content and re-extract only changed documents, then retract facts whose only evidence was the old version. Track freshness per source (last successful sync) and expose it, because a stale dependency map is worse than none during an outage. Periodic full reconciliation, comparing graph counts and keys with the sources, catches drift that incremental feeds miss. Our guide to getting enterprise knowledge ready for AI covers the ownership side of this.
Graph algorithms, embeddings and GNNs
22. Explain PageRank and a business use for it.
Answer: PageRank scores a node by the importance of the nodes that point to it, computed iteratively: each node shares its score across its outgoing edges, with a damping factor representing a chance of jumping to a random node. Highly scored nodes are those many important nodes depend on or point to. Business uses: identifying critical shared services in an IT dependency graph, influential accounts in a transaction network, or key parts in a bill-of-materials graph. Caveats: edge direction matters (point edges toward what is depended on), results depend on how you project the graph, and personalised PageRank, which restarts from a chosen set of seed nodes, is often more useful because it answers "what is important relative to this customer or incident".
23. What is community detection, and how do Louvain and Leiden differ?
Answer: Community detection partitions a graph into groups of nodes that are more densely connected to each other than to the rest. Louvain greedily optimises modularity by moving nodes between communities and then collapsing communities into super-nodes, repeating at each level. Leiden improves on Louvain by adding a refinement step that ensures communities are well connected internally; Louvain can produce communities that are internally disconnected. Leiden is the algorithm Microsoft's GraphRAG uses to build its community hierarchy. Other options are label propagation (fast, less stable) and weakly connected components, which is not community detection but is often the first step in fraud ring analysis.
24. How do you measure similarity between nodes?
Answer: Topological similarity compares neighbourhoods: Jaccard similarity is shared neighbours divided by total distinct neighbours, and overlap or Adamic-Adar weight rare shared neighbours more. These are explainable ("these two customers share three devices and an address"). Embedding similarity compares learned vectors (see Q25) and can capture structural roles that do not share any neighbours. Property similarity compares attributes. In practice, a node similarity or k-nearest-neighbour job writes SIMILAR edges with a score, which then feed recommendations, entity resolution candidates or "customers like this" features. Limit the top-k per node, or the similarity graph becomes larger than the original.
25. What are graph embeddings, and how are they produced?
Answer: A graph embedding is a vector per node (sometimes per edge or per graph) that encodes its position and connections, so nodes that are close or structurally alike get similar vectors. Random-walk methods such as DeepWalk and node2vec generate walks and train a word2vec-style model on them; node2vec's parameters bias walks toward local neighbourhoods or broader structure. Fast projection methods such as FastRP produce embeddings through random projections of the adjacency structure. Knowledge graph embedding models such as TransE, DistMult and ComplEx learn vectors for entities and relation types so that true triples score higher than false ones, which is useful for link prediction. These are different from text embeddings, explained in embeddings explained; many systems combine both.
26. Explain graph neural networks conceptually.
Answer: A graph neural network learns node representations by message passing: in each layer every node collects information from its neighbours, aggregates it (sum, mean, attention-weighted) and combines it with its own features through learned weights. After k layers a node's representation reflects its k-hop neighbourhood. Common families are GCN (normalised neighbour averaging), GraphSAGE (samples neighbours and learns aggregation, so it can generalise to unseen nodes, which matters for new customers), and GAT (attention over neighbours). Tasks are node classification (is this account a mule?), link prediction (will these two parts be used together?) and graph classification. Practical issues: oversmoothing with too many layers, neighbourhood explosion on high-degree nodes, and leakage when training edges include information from the future.
27. When would you use a GNN instead of simpler graph features with a gradient-boosted model?
Answer: Start simple. Hand-built graph features (degree, PageRank, community ID and size, shared-device counts, distance to a known fraud node) fed into a gradient-boosted model are fast, explainable and often strong. Move to a GNN when the signal lies in complex neighbourhood patterns that features miss, when you have enough labelled data, and when you can operate the training and inference pipeline, including point-in-time graph snapshots to avoid leakage. In regulated settings such as lending or fraud decisions, explainability requirements can favour the feature approach, or a GNN used to generate alerts that analysts then review.
GraphRAG and graphs versus vectors
28. What problem does GraphRAG solve that vector RAG does not?
Answer: Vector RAG retrieves the top-k chunks most similar to the question, which works when the answer sits in one or two passages (see what RAG is). It struggles with multi-hop questions where each link lives in a different document, with corpus-wide questions ("main themes across two years of complaints"), and with relationship questions ("how is this vendor connected to that outage?"). GraphRAG adds explicit entities and relationships so retrieval can follow connections and use pre-computed summaries. The full approach comparison is in our GraphRAG guide; in an interview, add that many apparent vector failures are really chunking or keyword problems that hybrid search and reranking fix more cheaply.
29. Explain local search versus global search in Microsoft's GraphRAG.
Answer: Local search is for questions about specific entities. It finds the entities that match the question, then gathers their neighbours, relationships, related source text and the summaries of communities they belong to, and fits that into the context. Global search is for questions about the whole corpus. It works map-reduce style over community summaries at a chosen level of the hierarchy: each summary produces a partial answer with a relevance score, and the highest-rated partial answers are combined. Global search is more expensive at query time and depends on summary quality; local search is cheaper and more precise but cannot answer "what are the themes?" The project has also described variants such as DRIFT search, which mixes community-level context into local search, and lower-cost approaches that defer summarisation to query time; check the current project documentation for details.
30. How are communities and community summaries built, and what are their weaknesses?
Answer: After extraction, the entity graph is partitioned with hierarchical Leiden into levels, from a few broad communities down to many small ones. An LLM then writes a report per community: key entities, relationships, claims and themes, built from the underlying entity and relationship descriptions. Weaknesses to raise: indexing cost grows with corpus size because every chunk and every community needs LLM calls; summaries are generated text and can contain errors or omissions that propagate into answers; a small change in the graph can reshuffle communities, so incremental updates are hard; and citations become indirect, because the answer cites a summary that cites entities that cite chunks. Keep the chain of IDs so you can still show the source passage.
31. When do graphs beat vectors, and when do vectors win?
Answer: Graphs win when the answer is defined by connections: multi-hop chains, paths, "everything affected by X", counts and filters over entities, ownership or dependency structures, and when exactness and explainability matter. Vectors win for fuzzy language over unstructured text, paraphrase, long-tail questions nobody modelled, and quick time to value with no schema. Most production systems use both: vector or hybrid search to find entry points, then a bounded graph traversal for connected facts, as our vector database interview questions discuss from the vector side.
| Question type | Better fit |
|---|---|
| "What does our travel policy say about international flights?" | Vector or hybrid search |
| "Which customers are two hops from a known mule account?" | Graph query |
| "Which services fail if this database goes down?" | Graph traversal |
| "Summarise complaints about a supplier and its subsidiaries" | Graph to scope, vectors for text |
| "Main themes in this year's audit findings" | Community summaries or a batch analysis |
32. How do you make text-to-Cypher reliable and safe?
Answer: Give the model a compact, accurate schema (labels, relationship types with direction, property names and example values) plus a few tested example queries. Prefer a library of parameterised query templates for frequent questions and use free-form generation as the fallback. Before execution: parse the query, reject write clauses (CREATE, MERGE, SET, DELETE, procedure calls that write), validate labels and types against the schema, enforce bounds on variable-length paths, a LIMIT and a timeout, and run under a read-only role with the user's permissions. If a query fails or returns nothing, feed the error back for one bounded retry. Evaluate on a set of question and expected-result pairs, comparing results rather than query strings. The same discipline appears in our text-to-SQL agent project.
33. How do you evaluate a GraphRAG or knowledge graph question-answering system?
Answer: Evaluate each layer separately. For the graph: precision and recall of extracted entities and relations, duplicate rate after resolution, and constraint violations. For retrieval: did the retrieved subgraph contain the facts needed (path recall), and was it small enough to fit the context. For text-to-Cypher: execution accuracy against expected results. For answers: faithfulness to the retrieved facts, correctness, and citation accuracy, using the same metrics family as vector RAG (see our RAG evaluation metrics). Build a test set by question type (single-hop, multi-hop, aggregation, global) and compare against a vector-only baseline, so you can show where the graph actually earns its cost.
If you want guided, hands-on practice with RAG pipelines, embeddings, evaluation and agents that pick between retrieval tools, Cloudsoft's APEX AI, ML, Cloud and Cyber Security program covers these skills through labs, in our Ameerpet classroom or live online.
Quality, provenance, security and scale
34. How do you model provenance and confidence in a knowledge graph?
Answer: Every fact should answer "who said so, when, and how sure are we?" In a property graph, put provenance on edges and nodes: source system, record ID or document and chunk ID, extraction method (master data, rule, model), model and prompt version, extraction run, timestamp, confidence, and review status. When several sources assert the same fact, either keep one edge with a list of sources or a separate Assertion or Evidence node per source linked to the fact, which lets you resolve conflicts and retract one source cleanly. In RDF, named graphs per source or batch play the same role. Queries can then filter by trust: a compliance report might use only master data and reviewed facts, while exploratory search includes model-extracted ones with a label.
35. What data quality checks would you run on a knowledge graph?
Answer: Run them as gates in the pipeline and as scheduled monitors:
- Constraints: uniqueness on business keys, required properties, allowed values (SHACL shapes in RDF, constraints plus validation queries in property graphs).
- Structural checks: edges with wrong endpoint types, orphan nodes, impossible cycles (a company that is its own parent), sudden degree spikes on one node.
- Duplicates: candidate pairs above a similarity threshold that were not merged, tracked as a rate over time.
- Freshness and completeness: last sync per source, counts compared with sources, share of nodes with key relationships (every service has an owner).
- Accuracy: a periodic sample of facts checked by humans against the source.
Publish these as a quality dashboard owned by a named team, otherwise nobody fixes them.
36. How do you enforce access control in a knowledge graph?
Answer: The difficulty is that graphs make it easy to traverse from data you may see to data you may not, or to infer hidden facts from structure (an edge to a "Whistleblower Case" node reveals something even if the case details are hidden). Approaches:
- Database-level role-based access: some engines offer fine-grained privileges on labels, relationship types and properties (Neo4j Enterprise is one example); others offer only database-level or IAM-based access, so check current documentation.
- Separate graphs or databases per sensitivity tier or tenant when the boundary is hard.
- Security labels on nodes and edges (classification, tenant, region) enforced by a query service that injects filters into every query, so the model or user never composes raw queries against the full graph.
- Policy for derived outputs: an embedding, a community summary or a PageRank score computed over restricted data can leak it, so compute them per access tier.
For AI agents, the query tool must run with the end user's identity, not a service account that sees everything; our guide to AI agent identity and access covers the pattern.
37. What is a supernode, and how do you deal with it?
Answer: A supernode is a node with a huge number of relationships: a country node linked to every customer, a "Microsoft Windows" node linked to every server, a popular merchant. Traversals through it explode, write locks on it cause contention, and it dominates algorithms like PageRank. Fixes: model the attribute as a property instead of a node if you never traverse through it; use more specific relationship types so traversals can skip it (LIVES_IN_CITY rather than a generic link); bucket edges by time or category (fan-out nodes per month); exclude known supernodes from algorithm projections; and stop traversals at nodes above a degree threshold. In fraud graphs, shared nodes like a bank's own branch address or a common payment gateway must be excluded or they connect everyone.
38. How do you scale a graph database?
Answer: First, scale reads with replicas and causal or eventual consistency settings appropriate to the use case. Many graphs fit on one large primary with read replicas; that is often simpler and faster than distribution. Partitioning (sharding) a graph is hard because traversals cross partitions; network hops replace memory hops, so partition along natural boundaries where few edges cross (tenant, region, business unit), or use an engine built for distribution. Separate operational queries (short, indexed lookups and bounded traversals) from analytics (full-graph algorithms), running analytics on an in-memory projection, a separate analytics engine or a batch export. Other levers: indexes on lookup properties, bounded queries, batched writes with UNWIND, avoiding hot nodes, and caching frequent subgraph results.
39. How would you load a hundred million relationships efficiently?
Answer: For an initial load, use the database's bulk import tool or bulk loader from files (for example Neptune's bulk loader from S3, or Neo4j's offline admin import), which bypasses transactional overhead. For incremental loads, batch rows (thousands per transaction) with UNWIND $rows AS row MERGE ..., create uniqueness constraints and indexes before merging, load nodes before relationships, and sort or partition input to reduce lock contention on shared nodes. Make every batch idempotent and record a load run ID so failed batches can be retried. Measure throughput on a sample, then scale; and plan how you would reload from scratch, because a full rebuild is sometimes the fastest way to fix a bad schema decision.
40. Design a knowledge graph platform for an enterprise that wants both analytics and AI assistants on top of it.
Answer: Separate ingestion, the curated graph, and consumers, with governance across all three:
sources: ERP, CRM, CMDB, HR, documents
| CDC / events / batch
v
ingest + map to ontology + entity resolution
| SHACL/constraint gates, provenance stamps
v
curated graph (operational) --> analytics projection
| (algorithms, GNN)
v
graph query service: templates, RBAC filters,
limits, audit log
|
+--> apps and dashboards
+--> AI assistant tools (text-to-Cypher,
GraphRAG retrieval with vector index)
Key decisions to explain: an ontology owned by a data governance group, stable IDs shared between the graph and the vector index so citations work across both, a single query service so access control and limits are not reimplemented in every app, algorithm results written back as properties on a schedule, and observability on query latency, freshness and quality metrics. Tie it to business questions and owners; a platform without a first use case rarely survives budget review. Our AI system design interview questions cover the surrounding LLM architecture.
Real-world scenario questions
41. Entity resolution wrongly merged two different customers with the same name and city. One customer's loans now appear on the other's profile. What do you do?
Answer: Contain first, then split, then fix the cause. Freeze the merged entity from downstream use (assistants, credit decisions, statements), split it back into two entities using the preserved source records, reassign each edge according to the source record it came from, and notify the teams who may have acted on the wrong view. Because a merged entity may have spread wrong facts (derived scores, communities, cached summaries), recompute those for both customers.
What I would check:
- Which match rule or score caused the merge, and whether strong identifiers (PAN, date of birth, customer ID) disagreed and were ignored.
- Whether source records were preserved, or overwritten during the merge.
- Every edge and derived property on the merged node, with its provenance, to reassign correctly.
- Other merges by the same rule version, to find similar errors.
- Who accessed the merged profile while it was wrong, for privacy and audit reporting.
Production consideration: Design for reversibility: keep source records as separate nodes linked to the resolved entity with RESOLVED_TO edges that carry the rule version and score, add negative constraints ("strong ID mismatch blocks auto-merge"), and keep a do-not-merge list that reviewers can update. In a bank this is also a data protection incident, so follow the incident process; see our guide to generative AI in banking for the regulatory context.
42. A user asks the assistant: "Which of our Tier-1 suppliers share a parent company with a vendor that had a quality incident last quarter?" Vector RAG gives a vague answer. How do you design this?
Answer: This is a multi-hop question with a time filter, so the answer should come from a graph query, with text used only to explain. Route it to a parameterised template or text-to-Cypher:
MATCH (s:Supplier {tier: 1})-[:SUBSIDIARY_OF*1..3]->(p)
<-[:SUBSIDIARY_OF*1..3]-(v:Supplier)
-[:INVOLVED_IN]->(i:QualityIncident)
WHERE s <> v AND i.date >= date($from)
RETURN s.name, p.name, v.name, i.id
LIMIT 200
Then fetch the incident reports by ID for a short explanation with citations.
What I would check:
- Whether ownership edges exist and are complete for Tier-1 suppliers (often they come from a third-party data feed).
- Whether incidents are linked to the resolved supplier entity, not to a name string.
- The depth bound and direction on SUBSIDIARY_OF, and whether "parent" means ultimate or immediate parent.
- How the router recognised this as a structured question.
Production consideration: Show the path in the answer ("Supplier A, owned by Group P, which also owns Vendor V, incident Q-1182") so users can verify it. A manufacturer's supplier-risk use case like this one is discussed in AI in manufacturing.
43. A core database cluster is failing. The incident commander asks, "What is affected?" How does a dependency graph help, and what can go wrong?
Answer: Traverse upstream from the failing component along dependency edges to list affected services, then to the business capabilities and customers they support, ranked by criticality:
MATCH (db:Component {id: $id})
<-[:DEPENDS_ON*1..6]-(svc:Service)
OPTIONAL MATCH (svc)-[:SUPPORTS]->(cap:Capability)
RETURN svc.name, svc.tier, collect(DISTINCT cap.name)
ORDER BY svc.tier
What I would check:
- Freshness of the CMDB and discovered dependencies; when was each edge last confirmed?
- Whether edges came from declared configuration only, or also from observed traffic (tracing, flow logs), which catches undocumented dependencies.
- Redundancy: a service with a failover path to another cluster may not be affected, so edges need a "hard or soft dependency" property.
- Owners and on-call contacts for each affected service.
Production consideration: A wrong dependency map during an incident misleads people under pressure, so show edge age and source in the output. After the incident, compare the predicted impact with what actually broke, and fix the edges that were missing.
44. Design a fraud ring detection approach using a graph for a bank's payments data.
Answer: Build a graph of customers, accounts, devices, phone numbers, emails, addresses, IP addresses and transactions. Rings show up as dense clusters of accounts sharing identifiers, and as money moving quickly through chains of new accounts. The method: exclude benign shared nodes (bank branch addresses, corporate gateways, carrier-grade NAT IPs); compute weakly connected components on shared-identifier edges; score components by size, account age, shared-device density and flow patterns; run community detection on transaction edges; then generate features (distance to known fraud, component size) for a model and an alert queue for investigators.
What I would check:
- Which shared identifiers are strong (device fingerprint) versus weak (a common surname).
- False positives from families, hostels and offices sharing addresses or Wi-Fi.
- Latency needs: real-time scoring at payment time versus nightly batch.
- Label quality: confirmed fraud cases with dates, to avoid training on the future.
Production consideration: Investigators need an explainable view: the subgraph and the reason codes, not just a score. Keep decisions to block or close accounts with humans, and log what the analyst saw.
45. A retailer wants a "customer 360" knowledge graph across its app, stores and call centre. How would you approach it?
Answer: The core is identity: resolve the customer across loyalty IDs, phone numbers, emails, device IDs and store transactions, with confidence levels and consent attached. Then connect orders, products, returns, service tickets, campaigns and preferences. Start with two or three concrete uses (service agents seeing full context, personalised offer features, churn signals) rather than "everything about the customer".
What I would check:
- Consent and purpose: which data may be linked for which purpose under the DPDP Act and the retailer's own policies.
- How household and shared identifiers are handled (one phone number used by several family members).
- Which system is the master for each attribute when sources conflict.
- How quickly app and store events must appear in the graph.
Production consideration: Make erasure work: a deletion request must remove the person's nodes, edges and derived artefacts (embeddings, summaries, features). See our guide on the DPDP Act for AI applications, and the sibling recommendation systems interview questions for how graph features feed recommenders.
46. A compliance team needs to map regulations to internal policies, controls and systems, and see gaps. How would you model it?
Answer: Model Regulation broken into Requirement nodes (clause-level), linked to Policy sections that address them, to Control nodes that implement policies, to System and Process nodes where controls operate, and to Evidence (test results, audit findings) with dates. Gap analysis becomes a query: requirements with no control path, controls with no recent evidence, systems in scope with no mapped controls. An LLM can suggest requirement-to-policy mappings from text, but each mapping stays "proposed" until a compliance owner approves it.
What I would check:
- Versioning: which version of a regulation or policy a mapping refers to, and what changes when a new circular arrives.
- Many-to-many mappings and partial coverage (a control may only partly satisfy a requirement).
- Ownership of each requirement and control.
Production consideration: Auditors will ask "who approved this mapping and when?", so approval metadata and history are part of the model, not an afterthought.
47. Graph queries that were fast last month now time out. What do you investigate?
Answer: Usually the data shape changed, not the code. Profile the slow queries and look for growth in specific nodes or relationship types.
What I would check:
PROFILEorEXPLAINplans: label scans instead of index seeks, or rows exploding at a variable-length step.- New supernodes: a new data feed may link many nodes to one shared value.
- Missing or dropped indexes and constraints after a migration.
- Unbounded patterns introduced by new templates or generated queries.
- Memory pressure: page cache too small for the grown store, causing disk reads.
- Concurrent bulk loads or algorithm jobs competing with interactive queries.
Production consideration: Add query timeouts, per-query row limits and a slow-query log with alerts, and run analytics on a separate projection or replica.
48. Your LLM extraction pipeline produced thousands of new relationship types like PARTNERED_WITH, PARTNERS_WITH and COLLABORATES_WITH. How do you fix it?
Answer: The extraction ran without a closed schema. Fix the source of the problem, then clean up. Define an allowed list of relationship types with definitions and examples, enforce it in structured output validation, and give the model an "other" path that goes to a review queue rather than creating a new type. For existing data, cluster the invented types (string similarity and embeddings of their names and descriptions), map each cluster to a canonical type or reject it with a human reviewing the mapping, then rewrite edges in batches, keeping the original type as a property for traceability.
What I would check:
- Whether some invented types reveal genuine gaps in the schema that the business needs.
- Direction and meaning: COLLABORATES_WITH may not mean the same as a contractual partnership.
- Downstream queries and text-to-Cypher prompts that referenced the old types.
Production consideration: Add a monitor on the count of distinct types and on new types per run, so schema drift is caught in one batch, not after months.
49. A text-to-Cypher assistant returned salary-band data to an employee who should not see it. How do you respond and redesign?
Answer: Treat it as a security incident: disable the tool or restrict it to safe templates, find what was exposed and to whom from the query audit log, and inform the security and HR data owners. The root cause is almost always that queries ran under a service account with full read access, relying on the prompt to "not query salary". Redesign so the database enforces it: run queries under the end user's identity or a role derived from it, apply property- and label-level restrictions or a query service that injects security filters, remove sensitive properties from the schema shown to the model, and block queries that touch restricted labels before execution.
What I would check:
- Exact queries generated and executed, and the role they ran under.
- Whether the restricted data came from a property, a related node, or a derived artefact such as a summary.
- Other users who could have triggered similar queries.
Production consideration: Add red-team tests for access boundaries to the evaluation set and run them on every prompt or schema change. The sibling AI security interview questions go deeper on this class of failure.
50. Leadership asks whether to build a knowledge graph for the company's internal assistant or just improve vector RAG. How do you decide?
Answer: Decide from the questions and the failure analysis, not from enthusiasm. Collect a few hundred real user questions, label them by type (single-document lookup, multi-hop, aggregation, relationship, corpus-wide), and analyse where the current system fails. If most failures are retrieval quality on single documents, fix chunking, hybrid search and reranking first. If a meaningful share of high-value questions need connections across systems, and structured sources already exist (CMDB, ERP, HR), build a targeted graph for that domain, starting from master data, and add it as a retrieval tool alongside vectors.
What I would check:
- Question mix and business value per question type.
- Availability and quality of structured sources and owners for the schema.
- Indexing and maintenance cost, including LLM extraction calls and human review.
- Skills to run a graph database in production.
Production consideration: Run a time-boxed pilot on one domain with an evaluation set that compares graph-plus-vector with vector-only, and agree in advance what improvement would justify the ongoing cost. This is the kind of evidence-based decision that forward deployed engineer interviews also probe.
Key takeaways
- A knowledge graph is a curated model of meaning and identity; the graph database is only the engine.
- Design the schema from real questions, use specific relationship types, and model events as nodes when they have their own facts.
- Entity resolution is the hardest and most consequential step: keep source records, prefer strong identifiers, and make merges reversible.
- LLM extraction needs a closed schema, structured output validation, evidence spans and provenance on every edge.
- Bound every traversal, watch for supernodes, and separate operational queries from analytics.
- Graphs beat vectors for connections, paths and exact filters; vectors win on fuzzy text; production systems usually combine them.
- Enforce access control in the database or query service, never in the prompt, and remember derived artefacts can leak data.
Interview preparation checklist
- Model one domain you know (a bank, a hospital, an IT estate) as a property graph on paper and explain each design choice.
- Write by hand: a two-hop Cypher match, a safe
MERGEload, a bounded variable-length path, a shortest path and a SPARQL query with a property path. - Load a public dataset into a local graph database and run PageRank, weakly connected components and Leiden or Louvain on it.
- Build a small entity resolution pipeline with blocking, scoring and a review threshold, and measure precision and recall.
- Build a mini GraphRAG pipeline with schema-constrained LLM extraction and compare it with vector-only retrieval on ten multi-hop questions.
- Prepare one story each on provenance, access control and a supernode or performance problem.
- Revise related topics in the sibling RAG interview questions and SQL interview questions for data and AI, since many panels mix them.
FAQ
What skills are required for a knowledge graph engineer role?
You need data modelling, at least one graph query language such as Cypher or SPARQL, Python for pipelines, entity resolution techniques, basic graph algorithms, and an understanding of data quality and governance. For GraphRAG roles, add LLM extraction, embeddings and RAG evaluation.
Should I learn Cypher or SPARQL first?
Learn Cypher first if you are targeting property graphs, GraphRAG or fraud and IT use cases, since its syntax also underpins the ISO GQL standard. Learn SPARQL if you are targeting life sciences, publishing, government data or teams that use RDF and ontologies.
Do I need to know Neo4j specifically for graph database interviews?
Not necessarily. Neo4j is common, so familiarity helps, but interviewers mainly test modelling, query patterns and trade-offs. If a role uses Amazon Neptune or another engine, learn its supported query languages and loading tools before the interview.
How should a fresher prepare for knowledge graph interview questions?
Learn the fundamentals in this guide, write queries by hand, and build one small project: load a dataset into a graph database, resolve duplicate entities, run a few algorithms and answer five multi-hop questions. Be ready to explain every modelling decision.
Are ontology interview questions only asked for semantic web roles?
They are most common in RDF and semantic web roles, but many GraphRAG and data platform interviews now ask about schemas, taxonomies versus ontologies, and how to stop LLM extraction from inventing relationship types.
Is GraphRAG replacing vector RAG?
No. GraphRAG adds value for multi-hop, relationship and corpus-wide questions, but vector and hybrid search remain the cheaper default for most document questions. Most production systems combine the two.
Is knowledge graph engineering a good career choice?
It can be a strong specialisation for engineers who enjoy data modelling and messy integration problems. Graph skills are used in fraud, supply chain, IT operations, compliance and AI retrieval work, and they combine well with data engineering and AI engineering roles.
How long does it take to prepare for a graph database interview?
It depends on your background. A developer comfortable with SQL and Python can usually cover the core topics and one project in a few focused weeks; a fresher should plan for longer and spend most of the time building and querying a real graph.
Knowledge graphs only pay off when they are connected to real systems, secured properly and evaluated against simpler options. To build those skills alongside RAG, agents, cloud and security, explore Cloudsoft's APEX program for AI, ML, cloud and cyber security. If you want to take retrieval systems into customer environments end to end, from discovery to deployment and evaluation, the AI Forward Deployed Engineer FDE PRO program covers that path, with placement support until you're placed. Classes run in Ameerpet, Hyderabad, or live online; call +91 96660 19191 for a free demo.



