New batches starting this week Β· Limited seats

Spring AI and Java AI Interview Questions and Answers 2026 (55 Questions)

55 Spring AI and Java AI interview questions with senior-level answers, covering ChatClient, advisors, structured output, tool calling, VectorStore RAG, MCP, LangChain4j, concurrency, testing, Kubernetes, security and real-world scenarios.

Spring AI and Java AI interview questions 2026: 55 questions on ChatClient, advisors, RAG, tools, VectorStore, MCP and LangChain4j
Last updated Β· 39 min read Β· 8,632 words

Spring AI interview questions now check whether a Java developer can put a large language model inside a real Spring Boot service: calling it through ChatClient, getting typed records back, letting it call @Tool methods safely, grounding it with a VectorStore, and running all of it with the same testing, observability and security discipline as any other production dependency. This guide covers 55 high-value Spring AI and Java AI interview questions, from fundamentals to scenarios, including LangChain4j, MCP, virtual threads, Kubernetes and migrating an existing service, with answers written the way a senior Java engineer would give them.

If you want the broader "why Java for AI" picture and a first walkthrough of the patterns, read AI for Java developers first. This page assumes that background and goes deeper into what interviewers probe. Code snippets are short and illustrative: both frameworks release often, so check class and method names against the version you use.

How to use this guide

  • Freshers and Java developers new to AI (Q1 to Q11): interviewers want clear mental models: ChatModel vs ChatClient, advisors, structured output, memory, and how Spring AI differs from LangChain4j. Explain them without buzzwords.
  • Working Spring Boot developers (Q12 to Q27): expect hands-on questions on tool calling, RAG with a VectorStore, ETL, chat memory storage and advisor ordering. Be ready to sketch code on a whiteboard.
  • Senior, lead and platform roles (Q28 to Q43): MCP, observability, concurrency, resilience, evaluation, testing, Kubernetes and security. These answers need trade-offs, not feature lists.
  • Everyone (Q44 to Q55): scenario questions decide most loops. Practise answering them aloud in a structured way: what you would check, in what order, and what you would change in production.

Contents

Spring AI fundamentals

1. What is Spring AI, and what problem does it solve for Java teams?

Answer: Spring AI is the Spring project for building AI features into Spring applications. It gives portable abstractions over model providers (chat, embeddings, images, audio), vector stores and tools, wired in through Spring Boot starters and auto-configuration. The problem it solves is integration, not model training: a Java team can call hosted or local models, ground answers on its own data and expose tools without leaving the Spring ecosystem or writing provider-specific HTTP clients.

The real value shows up in the "boring" parts: configuration properties, dependency injection, Micrometer observability, retries, and the same testing and deployment pipeline the team already uses.

Interview tip: Say clearly that Spring AI is an application framework, not a model or a training library. Candidates who confuse it with ML training tooling lose credibility early.

2. Which Spring AI version line goes with which Spring Boot version?

Answer: The current Spring AI 2.x line is built for Spring Boot 4 (the 2.0 documentation lists support for Boot 4.0.x and 4.1.x). The 1.x line (1.0 GA in 2025, then 1.1) targets Spring Boot 3, with 1.1 documented against Boot 3.4.x and 3.5.x. So a service still on Boot 3 uses Spring AI 1.1; moving to Spring AI 2 means moving to Boot 4 at the same time, along with Boot 4's own changes (for example the move to Jackson 3).

Always import the spring-ai-bom so every Spring AI module stays on one consistent version.

Interview tip: Interviewers like hearing that you check the compatibility table before upgrading, rather than bumping one dependency and hoping.

3. What is the difference between ChatModel and ChatClient?

Answer: ChatModel is the low-level, provider-facing interface: you hand it a Prompt and get a ChatResponse. ChatClient is the fluent, high-level API most application code should use. It adds default system prompts, prompt templates, advisors (memory, RAG, logging, safety), tool calling and structured output via entity().

In Spring AI 2.x this split matters more than before: ChatModel implementations no longer run an internal tool-execution loop. Calling chatModel.call(prompt) with tools returns the raw tool-call request; the loop is handled by a ToolCallingAdvisor that ChatClient registers automatically.

// Illustrative: Spring AI 2.x
@Service
class SummaryService {
  private final ChatClient chat;

  SummaryService(ChatClient.Builder builder) {
    this.chat = builder
        .defaultSystem("Summarise in plain English.")
        .build();
  }

  String summarise(String text) {
    return chat.prompt().user(text)
        .call().content();
  }
}

4. What is LangChain4j, and how is its programming model different?

Answer: LangChain4j is an independent open-source Java library for LLM applications. It works with plain Java, Spring Boot and Quarkus. Its signature feature is AI Services: you declare a Java interface, annotate methods with @SystemMessage and @UserMessage, and AiServices.builder(...) generates an implementation wired to a chat model, memory, tools and a content retriever. Spring AI leans on a fluent client plus advisors; LangChain4j leans on declarative interfaces.

// Illustrative: LangChain4j AI Service
interface PolicyAssistant {
  @SystemMessage("Answer only from the context.")
  String answer(@MemoryId String userId,
                @UserMessage String question);
}

PolicyAssistant bot = AiServices
    .builder(PolicyAssistant.class)
    .chatModel(chatModel)
    .chatMemoryProvider(id ->
        MessageWindowChatMemory.withMaxMessages(20))
    .build();

Interview tip: Current LangChain4j builders use chatModel(...); older tutorials show chatLanguageModel(...). Mentioning that you noticed the rename signals real hands-on use.

5. How do prompt templates and system prompts work in Spring AI?

Answer: You set a default system prompt on the builder with defaultSystem(...) and override or parameterise it per call with .system(s -> s.param("tone", tone)). User text can be a template too: .user(u -> u.text("Summarise {doc}").param("doc", text)). The default renderer is StTemplateRenderer (StringTemplate) with curly-brace delimiters. If your prompt contains JSON, switch delimiters (for example to angle brackets) so literal braces are not treated as variables.

Keep templates in versioned resource files, not string literals scattered across services, so prompt changes go through review and trigger your evaluation run.

6. How do you get structured output into a Java record?

Answer: Call .entity(MyRecord.class) on the call response. Spring AI uses a BeanOutputConverter to generate a JSON schema from the record, adds format instructions to the prompt and maps the reply back. For generic types use ParameterizedTypeReference, and use responseEntity(...) if you also need the ChatResponse metadata such as token usage.

record Triage(String category, String urgency,
              boolean needsHuman) {}

Triage t = chat.prompt()
    .user(complaint)
    .call()
    .entity(Triage.class);

Where the provider supports native structured output (schema-constrained JSON), Spring AI can use it instead of prompt-only instructions, which reduces malformed replies. Either way, validate the result: an enum-like field can still come back with an unexpected value. Our guide to function calling and structured outputs covers the model-side mechanics.

7. What is an EmbeddingModel and where does it fit?

Answer: EmbeddingModel turns text into a numeric vector that captures meaning, so similar texts have nearby vectors. In Spring AI you rarely call it directly in RAG: the VectorStore uses the configured embedding model when you add documents and when you search. You do call it directly for tasks such as deduplication, clustering or semantic caching.

Interview tip: Mention that the same embedding model must be used for ingestion and query, and that changing it means re-embedding the whole corpus. That is a migration, not a config tweak.

8. What is an advisor in Spring AI?

Answer: An advisor is an interceptor around a ChatClient call, similar in spirit to a servlet filter or AOP around-advice. It can modify the request before the model sees it and the response before your code sees it. Built-in advisors include MessageChatMemoryAdvisor and VectorStoreChatMemoryAdvisor (memory), QuestionAnswerAdvisor and RetrievalAugmentationAdvisor (RAG), SafeGuardAdvisor (blocks configured sensitive words), SimpleLoggerAdvisor (logging) and the auto-registered ToolCallingAdvisor. You register them as defaults on the builder or per call with .advisors(...).

9. How does chat memory work, and what changed in Spring AI 2?

Answer: An LLM is stateless, so "memory" means your application stores earlier messages and sends the relevant ones again. In Spring AI, ChatMemory decides what to keep. The default MessageWindowChatMemory keeps a sliding window (20 messages by default), always preserves system messages and evicts whole turns. Storage is a ChatMemoryRepository: in-memory by default, or JDBC, Cassandra, Neo4j, MongoDB, Redis and others.

chat.prompt()
    .user(question)
    .advisors(a -> a.param(
        ChatMemory.CONVERSATION_ID, sessionId))
    .call().content();

In 2.x the conversation ID is mandatory: the old default ID was removed, so a missing ID fails fast instead of silently mixing users into one shared conversation. PromptChatMemoryAdvisor was also removed in favour of MessageChatMemoryAdvisor.

10. What is the difference between call() and stream()?

Answer: call() blocks until the full reply is ready and returns content(), chatResponse() or entity(). stream() returns a Reactor Flux of tokens or partial responses, so the UI can show text as it is generated. Streaming improves perceived latency for chat; it does not make the model faster. Spring AI's notes say non-streaming calls rely on the servlet stack and streaming on the reactive stack, so a Spring MVC app that streams also needs WebFlux on the classpath.

Interview tip: Point out that structured output via entity() needs the complete reply, so it is a call() pattern.

11. Spring AI or LangChain4j: how do you choose?

Answer: Both are capable; the choice is mostly about your estate and team style.

FactorSpring AILangChain4j
Natural homeSpring Boot servicesQuarkus, plain Java, mixed estates
Main styleFluent ChatClient plus advisorsDeclarative AI Service interfaces
ConfigurationBoot properties and auto-configBuilders, plus Spring/Quarkus integrations
ObservabilityMicrometer observations built inListeners and integration-specific support
MCPClient and server startersStrong client; server side via integrations

A good answer ends with a decision rule: in a Spring-standardised GCC team, start with Spring AI; if half the services are Quarkus, LangChain4j gives one library across both. Keep the AI calls behind your own interface so the decision is reversible.

Tool calling

12. How does @Tool calling work in Spring AI from end to end?

Answer: You annotate a method with @Tool and a description, and parameters with @ToolParam. When you pass the object with .tools(...) (or defaultTools(...)), Spring AI generates a tool definition with a JSON schema and sends it with the prompt. If the model decides to call the tool, it returns a tool-call request instead of text; the ToolCallingAdvisor executes the method, sends the result back, and repeats until the model produces a final answer.

class OrderTools {
  @Tool(description = "Delivery status of an "
      + "order owned by the current customer")
  String status(
      @ToolParam(description = "Order id")
      String orderId) {
    return orders.statusFor(orderId);
  }
}

The model never runs your code; it only asks. Your code decides whether and how to execute. In 2.x only explicitly attached tools can run by default, and the old bean-name lookup (toolNames()) was removed.

Interview tip: The description is the model's only guidance on when to use a tool. A vague description causes wrong tool choices more often than bad code does.

13. What is ToolContext, and why should identity never come from model arguments?

Answer: ToolContext lets you pass data to a tool method that is not sent to the model, such as tenant ID, user ID or request scope: .toolContext(Map.of("tenantId", tenant)). That matters because anything the model fills in can be manipulated by prompt injection. If a "get account balance" tool accepts customerId from the model, a crafted message can make it ask for someone else's account.

The rule: take identity from Spring Security or ToolContext, use model arguments only for the things the user is allowed to choose, and authorise inside the tool as if the call came from an untrusted client.

Real-world example: Consider a bank's card-services assistant. The blockCard tool reads the customer from the security context, accepts only a card reference that is checked against that customer's cards, and returns a confirmation request rather than blocking immediately.

14. When would you take over the tool-execution loop yourself?

Answer: When you need control between steps: human approval before a write action, custom auditing, a step limit, or routing some tool calls to a different executor. In Spring AI 2.x you disable the auto-registered ToolCallingAdvisor for that call, receive the raw tool-call request, and drive the loop with a ToolCallingManager. (In 1.x the equivalent was the internalToolExecutionEnabled option, which 2.x removed.)

For most read-only tools the automatic loop is fine. For anything that moves money, changes records or messages customers, an explicit loop with an approval checkpoint is easier to defend in an audit.

15. How does tool calling work in LangChain4j, and what controls does it give you?

Answer: You annotate methods with @Tool and parameters with @P, then pass the objects to the AI Service builder with .tools(...), or supply a ToolProvider to choose tools dynamically per request. Useful controls include a tool execution error handler (send the error to the model, or fail the invocation), a strategy for hallucinated tool names, concurrent execution of multiple tool calls, and a return behaviour that sends a tool result straight back to the caller instead of through the model. @ToolMemoryId lets a tool see which conversation it is serving.

Interview tip: Mention a cap on sequential tool invocations. An agent stuck in a tool loop is a cost and latency incident.

RAG, VectorStore and ETL

16. What is the VectorStore abstraction in Spring AI?

Answer: VectorStore is one interface over many vector databases: add(List<Document>) embeds and stores documents, delete(...) removes by ID or filter, and similaritySearch(SearchRequest) retrieves. Implementations include PGvector, Elasticsearch, OpenSearch, Redis, MongoDB Atlas, Milvus, Qdrant, Pinecone, Weaviate, Neo4j, Oracle, Cassandra and others. SimpleVectorStore is for tests and demos only.

List<Document> hits = vectorStore
    .similaritySearch(SearchRequest.builder()
        .query(question)
        .topK(5)
        .similarityThreshold(0.6)
        .build());

The default topK is small (four), and the default similarity threshold accepts everything, so tune both against your evaluation set. If your team already runs PostgreSQL, PGvector is often the simplest start; the pgvector RAG tutorial walks through schema, indexing and queries.

17. How do you use metadata filters, and why do they matter for security?

Answer: Each Document carries metadata. SearchRequest accepts a portable filter expression, either as a string such as "dept == 'claims' && year >= 2025" or built with FilterExpressionBuilder. Spring AI translates it into the store's native query.

Filters are how you enforce document-level access in RAG. Tag chunks at ingestion with tenant, department and classification, and build the filter from the authenticated user on every request. Never let the model or the user's text choose the filter. With QuestionAnswerAdvisor you can pass a per-request filter through the FILTER_EXPRESSION advisor parameter.

18. QuestionAnswerAdvisor or RetrievalAugmentationAdvisor?

Answer: QuestionAnswerAdvisor is simple "naive" RAG: search the vector store with the user's question, append the chunks to the prompt, done. It is the right first version. RetrievalAugmentationAdvisor is a modular pipeline: query transformation and expansion before retrieval, one or more document retrievers, joining and post-processing of results, then query augmentation. Use it when you need rewriting of follow-up questions, multiple sources, or explicit control over what happens when nothing relevant is found.

Advisor rag = RetrievalAugmentationAdvisor
    .builder()
    .documentRetriever(VectorStoreDocumentRetriever
        .builder().vectorStore(store)
        .topK(5).build())
    .queryAugmenter(ContextualQueryAugmenter
        .builder().allowEmptyContext(false)
        .build())
    .build();

19. Which modular RAG components should a Java engineer know?

Answer: Pre-retrieval: RewriteQueryTransformer (cleans up vague queries), CompressionQueryTransformer (turns a follow-up plus chat history into a standalone query), TranslationQueryTransformer (translates queries into the embedding model's language) and MultiQueryExpander (several phrasings of one question). Retrieval: VectorStoreDocumentRetriever with threshold, topK and filters. Generation: ContextualQueryAugmenter, which by default refuses to answer when retrieval returns nothing, rather than letting the model improvise.

Real-world example: For a Hyderabad GCC's internal HR assistant, employees write in a mix of English and Telugu or Hindi. A translation transformer before retrieval over English policy documents improves recall without re-embedding the corpus in several languages. Each transformer is an extra model call, so measure the latency cost.

20. What does the Spring AI ETL pipeline do?

Answer: It is the ingestion side of RAG, modelled as three functional interfaces: DocumentReader (a supplier of documents), DocumentTransformer (a function over document lists) and DocumentWriter (a consumer, usually the VectorStore). Readers include PagePdfDocumentReader, ParagraphPdfDocumentReader, TikaDocumentReader (DOCX, PPTX, HTML and more), JsonReader, TextReader, MarkdownDocumentReader and JsoupDocumentReader. Transformers include TokenTextSplitter, KeywordMetadataEnricher and SummaryMetadataEnricher.

var docs = new TikaDocumentReader(resource).read();
var chunks = new TokenTextSplitter().apply(docs);
chunks.forEach(d -> d.getMetadata()
    .put("dept", "claims"));
vectorStore.add(chunks);

Run ingestion as a separate job (a scheduled task, Spring Batch, or a Kubernetes CronJob), not inside the request path. The RAG chunking strategies guide covers how chunk size affects answer quality.

21. How do you handle large ingestion jobs without hitting embedding limits?

Answer: Spring AI's VectorStore.add uses a BatchingStrategy (default TokenCountBatchingStrategy) to split documents into batches that fit the embedding model's input limit with a reserve. On top of that, design the job for production: process sources incrementally using a content hash so unchanged documents are not re-embedded, use stable document IDs so updates replace old chunks instead of duplicating them, delete chunks for removed source documents, and record per-run counts so a silent failure is visible.

Interview tip: "How do you keep the index in sync with SharePoint or Confluence?" is a common follow-up. Talk about change detection, deletes and re-indexing, not just the first load.

22. How is RAG built in LangChain4j?

Answer: Ingestion uses an EmbeddingStoreIngestor with a DocumentSplitter and an EmbeddingModel, writing into an EmbeddingStore such as PgVectorEmbeddingStore. Retrieval uses an EmbeddingStoreContentRetriever with maxResults, minScore and a metadata filter, attached to the AI Service with .contentRetriever(...). For advanced flows, a DefaultRetrievalAugmentor combines a QueryTransformer, a QueryRouter across several retrievers, a ContentAggregator (including re-ranking) and a ContentInjector. "Easy RAG" is a quick-start module for demos, not a production design.

The concepts map almost one-to-one onto Spring AI's modular RAG, which is a good point to make in an interview. For RAG theory beyond Java, see the RAG interview questions guide.

Chat memory, advisors and output

23. Where should chat memory live in production?

Answer: Not in the default in-memory repository once you run more than one pod: a user routed to a different replica loses context, and a restart wipes everything. Use a shared ChatMemoryRepository such as JDBC (PostgreSQL), Redis or Cassandra. Check the trade-offs: the docs note that some repositories, including JDBC and MongoDB, do not store tool-call messages, which matters for agent-style conversations. Upgrading the JDBC repository to Spring AI 2 also needs a schema migration (a new ordering column).

Add retention rules too: conversation history is personal data under laws such as India's DPDP Act, so define how long you keep it and how a user's history is deleted.

24. How does advisor ordering work, and why does it matter?

Answer: Each advisor has an order from getOrder(). Lower values run first on the way in and last on the way out, so the chain behaves like a stack. Order changes behaviour: memory should be added before RAG rewrites the query, a safety check should see the final prompt, and logging placed outermost sees everything. In Spring AI 2, the default memory precedence was moved so memory advisors sit outside the ToolCallingAdvisor, meaning one memory write per user turn rather than per tool iteration.

Interview tip: If you have hit a bug caused by advisor order, describe it. It is the kind of detail only hands-on developers know.

25. How would you write a custom advisor?

Answer: Implement CallAdvisor (and StreamAdvisor if you support streaming), give it a name and an order, and in adviseCall modify the ChatClientRequest, call chain.nextCall(request), then inspect the ChatClientResponse. Typical custom advisors: PII redaction before the prompt leaves your network, tenant-specific system instructions, a cost guard that rejects oversized prompts, or an output check that flags answers without citations.

class RedactionAdvisor implements CallAdvisor {
  public ChatClientResponse adviseCall(
      ChatClientRequest req,
      CallAdvisorChain chain) {
    return chain.nextCall(redactor.apply(req));
  }
  public String getName() { return "redact"; }
  public int getOrder() { return 0; }
}

Keep advisors small and single-purpose; a "god advisor" is as hard to test as a god class.

26. What do you do when structured output fails to parse?

Answer: Treat model output as untrusted input. First reduce the failure rate: use provider-native structured output where available, keep the record small with clear field names, use enums for categories, and put a @JsonPropertyDescription or field descriptions on ambiguous fields. Then handle what remains: catch the conversion exception, retry once with the error message, and fall back to a safe default such as "route to human". Validate business rules with Bean Validation after parsing; valid JSON can still contain an impossible value.

Log the failure rate as a metric. A sudden jump after a model upgrade is an early warning.

27. How do you use more than one model in the same Spring Boot application?

Answer: Create separate ChatClient beans, each built from its own ChatModel, and qualify them by purpose (for example classifierClient and reasoningClient). Spring AI's spring.ai.model.chat property selects which provider is auto-configured; for multiple providers you define beans explicitly, using the documented builder configurer so observability stays wired. Common reasons: a small, cheap model for classification and routing, a stronger model for complex answers, and a second provider as fallback.

Route by task, not by habit, and record the model used on every response for later analysis. See LLM latency optimization for how routing affects response time.

Building these Spring Boot fundamentals properly (REST APIs, Spring Security, JPA, testing and deployment) is what makes AI features land safely. Cloudsoft's Java Full Stack course covers that core Java and Spring Boot foundation, classroom in Ameerpet or live online.

MCP in Java

28. How does a Spring Boot service act as an MCP client?

Answer: MCP (Model Context Protocol) is an open protocol, introduced by Anthropic in late 2024, for connecting AI applications to tools and data. With the spring-ai-starter-mcp-client starter (or the WebFlux variant), you configure connections to MCP servers through properties, using STDIO for local processes or Streamable HTTP for remote servers. Spring AI exposes the servers' tools as ToolCallbacks that you attach to a ChatClient, so the model sees them like any local @Tool.

The production questions are about trust: which servers are allowed, which of their tools are exposed (filter, do not take all), how the client authenticates, and what happens when a server's tool descriptions change. The MCP interview questions guide goes deep on protocol and security details.

29. How do you expose an existing Spring service as an MCP server?

Answer: Add an MCP server starter: spring-ai-starter-mcp-server for STDIO, or spring-ai-starter-mcp-server-webmvc / -webflux for HTTP. The protocol property selects SSE, Streamable HTTP or stateless Streamable HTTP. Annotate methods with @McpTool (and @McpResource or @McpPrompt where useful), and the starter registers them.

@Component
class ClaimTools {
  @McpTool(description = "Status of a claim "
      + "the caller is allowed to view")
  ClaimStatus claimStatus(String claimId) {
    return claims.statusFor(claimId);
  }
}

For a horizontally scaled service on Kubernetes, prefer the stateless HTTP mode; it fits the MCP specification revision dated 2026-07-28, which made the protocol stateless (see MCP 2026 spec changes). Secure the endpoint with Spring Security like any API, and remember that the newer MCP Java SDK validates tool arguments against the schema by default.

30. How does LangChain4j handle MCP?

Answer: LangChain4j provides an MCP client: create a transport (StdioMcpTransport for a subprocess, StreamableHttpMcpTransport for remote servers, plus Docker and WebSocket options), wrap it in a DefaultMcpClient, and pass an McpToolProvider to the AI Service with .toolProvider(...). The provider can combine several clients, filter tools by name or predicate, and rename clashing tool names. Its documentation states support for both the earlier handshake-based protocol revision and the stateless 2026-07-28 revision with version detection.

Interview tip: Always mention tool filtering. Exposing every tool from a broad server to the model widens the attack surface and confuses tool selection.

Concurrency, streaming and resilience

31. How do virtual threads help with LLM calls in Spring Boot?

Answer: LLM calls are slow, I/O-bound and mostly waiting. On a classic platform-thread pool, each in-flight call holds a thread for seconds, so a modest burst of traffic can exhaust the Tomcat pool and starve unrelated endpoints. With Java 21 virtual threads (spring.threads.virtual.enabled=true in Spring Boot), a blocked call parks a cheap virtual thread instead, so you can keep the simple imperative call() style and still handle many concurrent requests.

Caveats worth saying aloud: virtual threads raise concurrency, not provider capacity, so you still need limits (a semaphore or bulkhead) matching your rate limits; pinning on synchronized blocks around I/O can hurt on older JDKs; and connection pools to databases still cap throughput.

// Fan out independent model calls
try (var ex = Executors
        .newVirtualThreadPerTaskExecutor()) {
  var a = ex.submit(() -> classify(text));
  var b = ex.submit(() -> extract(text));
  return merge(a.get(), b.get());
}

32. How do you build a streaming chat endpoint, and what breaks it?

Answer: Return the Flux<String> from .stream().content() as Server-Sent Events (produces = TEXT_EVENT_STREAM_VALUE). What breaks it in practice is the infrastructure, not the code: proxies or ingress controllers that buffer responses, idle timeouts shorter than a long answer, gzip buffering, and clients that do not reconnect. Also remember that tool calling is imperative and blocking, so a streaming flow with tools mixes blocking work into a reactive pipeline; keep it off the event-loop threads.

Handle cancellation: if the user closes the tab, the subscription should cancel so you stop paying for tokens nobody reads.

33. How do you make LLM calls resilient?

Answer: Layer it. Spring AI has built-in retry for transient provider errors (spring.ai.retry.* properties: attempts, exponential backoff, which status codes to retry). Check the defaults; they can be generous for a user-facing request, so tighten them per use case. Add timeouts on the HTTP client, a circuit breaker and bulkhead (for example Resilience4j) per provider, and a fallback: a second provider or region, a cached answer, or a graceful "I can't answer right now, here is the self-service link".

Do not retry non-idempotent tool actions blindly. If the model asked to raise a ticket and the response timed out, check whether the ticket exists before raising another.

34. How do you control cost and token usage?

Answer: Measure first: Spring AI's Micrometer metrics record token usage by type (input, output, total), and chatResponse() metadata exposes usage per call. Then reduce: trim prompts and RAG context (fewer, better chunks), cap output tokens, use a smaller model for simple tasks, cache answers to repeated questions where data allows, and avoid re-sending long histories by windowing memory. Attribute cost per feature and per tenant by tagging observations, so a finance team can see which feature spends what.

Real-world example: Consider a retailer whose product-question bot sends the whole catalogue description plus twenty history messages every turn. Halving retrieved chunks and windowing memory cuts input tokens with no quality loss on the evaluation set; the metric shows it the same day.

Observability, evaluation and testing

35. What does Spring AI give you for observability?

Answer: Micrometer observations for ChatClient and its advisors, ChatModel, EmbeddingModel, ImageModel, VectorStore and tool calling. That yields metrics (operation timings, active calls, token usage) and traces that export through Spring Boot's usual Micrometer and OpenTelemetry setup to Prometheus, Grafana, Jaeger or a vendor backend. In 2.x, tool spans are named execute_tool <tool-name>.

Prompt and completion content is not logged by default. Properties such as spring.ai.chat.client.observations.log-prompt enable it, and the docs warn that this can expose sensitive data. In regulated environments keep it off in production, or send content only to a restricted store with redaction. Our AI observability guide covers what to monitor beyond latency.

36. How do you trace one AI request end to end?

Answer: Let the trace context flow from the HTTP request through the ChatClient call, each advisor, the vector store query, every tool execution and any downstream REST or MCP call. With Micrometer Tracing and an OpenTelemetry exporter this happens largely automatically inside one service, and propagates to others via headers. Add business tags you will need during an incident: feature name, tenant, model, prompt version, retrieved document IDs (not their content).

HTTP POST /claims/summary
  ChatClient call
    advisor: memory
    advisor: retrieval
      vector store query (pgvector)
    model call (input/output tokens)
    execute_tool policyLookup
      REST GET policy-service

37. How do you evaluate answer quality in a Java project?

Answer: Keep a versioned evaluation set of realistic questions with expected outcomes and run it whenever prompts, models, retrieval settings or data change. Spring AI ships an Evaluator interface with RelevancyEvaluator (is the answer relevant to the question and context?) and FactCheckingEvaluator (is the claim supported by the context?), both using a model as judge. You call them from JUnit, typically in a separate test profile or pipeline stage because they cost money and take time.

var eval = new RelevancyEvaluator(
    ChatClient.builder(judgeModel));
var res = eval.evaluate(new EvaluationRequest(
    question, retrievedDocs, answer));
assertThat(res.isPass()).isTrue();

Judge scores are a signal, not truth; spot-check them by hand and track pass rates over time. The LLM evaluation guide covers datasets and metrics, and the AI testing interview questions cover the QA view.

38. How do you unit test code that uses ChatClient without calling a real model?

Answer: Split deterministic logic from the model. For unit tests, mock ChatModel with Mockito and build a real ChatClient on it, so your advisors, prompt assembly, output parsing and fallbacks run for real while the model returns a canned ChatResponse. Test the cases that matter: malformed JSON, an empty retrieval result, a tool-call request for a forbidden action, a timeout. For integration tests, use Testcontainers for PostgreSQL with pgvector, and either a local model container or a stubbed HTTP provider (WireMock) for repeatable responses.

ChatModel model = mock(ChatModel.class);
when(model.call(any(Prompt.class)))
    .thenReturn(new ChatResponse(List.of(
        new Generation(new AssistantMessage(
            "{\"category\":\"FRAUD\"}")))));
var svc = new TriageService(
    ChatClient.builder(model));

Interview tip: Say that unit tests prove your code is correct and evaluations prove the system is good enough. Interviewers want to hear both layers.

Architecture, Kubernetes and security

39. Sketch the architecture of an AI feature inside a Spring Boot microservice.

Answer: Keep the AI behind a normal service boundary so callers do not care that a model is involved.

Client -> API gateway -> Spring Boot service
              |
       Spring Security (user, tenant)
              |
        Domain service --> ChatClient
              |              | advisors:
              |              | memory, RAG,
              |              | redaction
              |              v
              |        model provider
              |
   @Tool methods -> core APIs / MCP servers
   VectorStore   -> PostgreSQL + pgvector
   Micrometer    -> metrics and traces

Key decisions to explain: where identity comes from (Security, not the prompt), what the model can call (an allow-list of tools), where data goes (provider region, DPDP and sector rules), how outputs are validated, and what the fallback is. The same structure works whether the model is Amazon Bedrock, Azure OpenAI, Gemini or a self-hosted model; Spring AI's abstraction makes provider changes mostly configuration.

40. What changes when you deploy a Spring AI service to Kubernetes?

Answer: Most of the usual Spring Boot on Kubernetes practice applies, with a few AI-specific points:

  • Secrets and identity: provider keys from a secret manager via External Secrets or the CSI driver, or better, workload identity (EKS Pod Identity or IRSA for Bedrock, workload identity for Azure and Google Cloud) so there is no static key at all.
  • Probes: readiness should not call the model on every probe; it costs money and couples your pod health to the provider.
  • Scaling: CPU is a poor signal for I/O-bound AI services. Scale on in-flight requests or latency, and cap replicas to stay inside provider rate limits.
  • Networking: egress policies to only the provider endpoints, and ingress timeouts and buffering settings that suit streaming.
  • Jobs: ingestion as a CronJob or separate deployment, not in the API pods.

The Kubernetes for FDE engineers guide covers these patterns in more depth.

41. How do you manage secrets and data exposure for model providers?

Answer: Never commit keys to application.yml; reference environment variables or a secret store, rotate keys, and use separate keys per environment and per service so you can revoke one without an outage. Prefer cloud IAM roles over API keys where the provider supports it. Beyond credentials, control data: redact PII before prompts leave your network where the use case allows, confirm the provider's data retention and region terms, keep prompt and completion logging off unless there is a reviewed reason, and treat the vector store as a copy of sensitive documents with the same access controls as the source.

42. How do you defend a Java AI service against prompt injection?

Answer: Assume injection will sometimes succeed, and limit what it can achieve. Prompt injection arrives directly in user text or indirectly through retrieved documents, emails or tool results. Defences, in order of value:

  1. Least-privilege tools: read-only by default, identity from the security context, authorisation inside each tool.
  2. Human approval for consequential actions (payments, record changes, external messages).
  3. Output validation: structured output with strict types, allow-listed actions, no executing model-generated SQL or code without checks.
  4. Input and context screening: separate trusted instructions from untrusted content in the prompt, and use a classifier or guardrail service for known attack patterns.
  5. Monitoring and red-teaming: log tool calls and refusals, and test with injection payloads in CI.

Spring AI's SafeGuardAdvisor blocks configured sensitive words; it is a useful extra layer, not an injection defence on its own. See AI guardrails for the wider control set.

43. How would you design a multi-tenant AI feature in a Spring SaaS product?

Answer: Carry the tenant from authentication through everything: a tenant-aware filter on every vector search, tenant ID in ToolContext for every tool, conversation IDs namespaced by tenant, and per-tenant rate limits and token budgets. Decide isolation level per data class: shared vector table with mandatory filters for low-risk content, separate schemas or stores for regulated tenants. Add tests that try to retrieve another tenant's document, and run them in CI. Tag metrics by tenant so one heavy customer cannot hide inside an average.

Scenario-based questions

44. You need to add an AI feature to an existing Spring Boot 3 service. How do you plan it?

Answer: Start from the business outcome, not the framework. Agree what the feature should do, how success is measured and what it must never do, then add the smallest version behind a feature flag.

What I would check:

  1. The current Boot version: stay on Spring AI 1.1 if the service is on Boot 3 and a Boot 4 upgrade is not planned yet; do not couple two risky upgrades.
  2. Data: which documents or APIs the feature needs, who may see them, and whether they can leave the network.
  3. Provider: an approved model and region, and how the service will authenticate.
  4. Integration seam: a new domain service wrapping ChatClient, so controllers and the rest of the code do not depend on Spring AI types.
  5. Evaluation set: twenty to fifty realistic examples agreed with the business before writing prompts.
  6. Observability, timeouts and a fallback path before the first user sees it.

Production consideration: Release to internal users first, then a small group, comparing evaluation and real feedback at each step. This is the "from AI demo to enterprise outcome" path: the demo takes a day, the controls take the rest of the sprint.

45. Your team must upgrade from Spring AI 1.x to 2.x. What will break?

Answer: The upgrade comes with Spring Boot 4, so plan it as a platform upgrade. Spring AI-specific breaks to look for include: ChatModel no longer running tools by itself (move tool code to ChatClient); internalToolExecutionEnabled and bean-name tool resolution removed; conversation ID now mandatory for memory advisors; PromptChatMemoryAdvisor removed; immutable options classes with builders; flattened configuration properties (the .options segment removed); a JDBC chat memory schema change; and MCP annotation packages moved into Spring AI.

What I would check:

  1. The official upgrade notes and any OpenRewrite recipes offered.
  2. Every place code calls ChatModel directly with tools.
  3. Configuration files for renamed properties (silent misconfiguration is worse than a compile error).
  4. Database migrations for chat memory tables.
  5. The full evaluation suite before and after, on the same model.

Production consideration: Upgrade one service first, run it alongside the old version on shadow traffic if possible, and compare outputs and metrics.

46. A complaint-triage endpoint intermittently throws parsing errors. How do you fix it?

Answer: Consider an insurer whose Spring Boot triage service uses .entity(Triage.class). A few requests a day fail conversion, and those complaints never get routed.

What I would check:

  1. The raw replies of failed calls (from a restricted log): truncated output, extra prose, or a wrong enum value?
  2. Whether max output tokens is too low for long complaints, causing truncation.
  3. Whether provider-native structured output is available and enabled.
  4. Whether the record has ambiguous or optional fields that the model fills inconsistently.
  5. Whether failures correlate with a model version change or very long inputs.

Production consideration: Whatever the root cause, a parsing failure must never drop a complaint. Catch it, retry once, then route to a human queue with a reason code, and alert on the failure-rate metric.

47. Users see answers based on documents from another department. What happened?

Answer: This is an access-control incident, not a quality bug. Treat it that way: contain first, then investigate.

What I would check:

  1. Whether every retrieval path applies the department filter, including new code paths (a second advisor, a query expander running searches without the filter).
  2. Whether the filter is built from the authenticated user, not from request text or a cached value.
  3. Whether ingestion tagged all chunks correctly; untagged chunks may pass a filter that only excludes known values.
  4. Whether chat memory or a cache is shared across users with different access.

Production consideration: Make filters default-deny (match only explicitly allowed values), add a CI test per role that attempts cross-department retrieval, and follow your incident process, since leaked content may need to be reported.

48. Under load, unrelated endpoints of an AI-enabled service start timing out. Why?

Answer: The classic cause is thread starvation: blocking LLM calls hold servlet worker threads for seconds, so health checks and simple CRUD calls queue behind them.

What I would check:

  1. Tomcat thread pool usage and queue depth during the incident.
  2. Model latency percentiles and provider throttling responses at the same time.
  3. Retry settings: aggressive retries multiply held threads during provider slowness.
  4. Database connection pool saturation from chat memory or vector queries.

Production consideration: Enable virtual threads, put a bulkhead around AI calls sized to provider limits, tighten timeouts and retries for user-facing requests, and consider moving the AI feature to its own deployment so it cannot starve core APIs.

49. A streaming answer works locally but arrives all at once in production. How do you debug it?

Answer: Something between the app and the browser is buffering.

What I would check:

  1. The response content type and that the controller returns the Flux as Server-Sent Events.
  2. Ingress or reverse-proxy buffering settings (for example NGINX proxy buffering) and response compression.
  3. API gateway or load balancer idle timeouts that also cut long streams.
  4. Any servlet filter in the app that wraps and buffers the response body.

Production consideration: Add an end-to-end synthetic test that measures time to first token through the real ingress, so a config change cannot silently break streaming again.

50. An audit shows a tool was called with another customer's account number. What do you change?

Answer: The tool trusted a model-supplied identifier. The fix is in the tool design, not the prompt.

What I would check:

  1. Which tool parameters identify a person or account, and where their values come from.
  2. Whether the tool re-authorises the caller against the resource, or relies on the downstream API.
  3. The conversation that triggered it: user error, injection through a pasted document, or a model mistake.

Production consideration: Remove identity parameters from tools and take them from ToolContext or Spring Security, authorise every call against the resource owner, and add an audit event per tool call with user, tool and arguments. The AI agent developer interview questions cover more agent safety patterns.

51. You find full prompts, including customer PII, in your log platform. What now?

Answer: Someone enabled content logging (for example log-prompt or a logging advisor at debug level) in production.

What I would check:

  1. Which properties or advisors emit content, in which environments, since when.
  2. Where the logs went, who can access them and their retention.
  3. Whether traces or span attributes also carry content.

Production consideration: Turn content logging off, purge or restrict the affected logs according to your incident and privacy process, add a startup check that fails if content logging is on in the production profile, and if content capture is needed for debugging, route it to a restricted store with redaction and short retention.

52. Your model provider has a regional outage during business hours. How should the service behave?

Answer: It should degrade, not fail. The design work happens before the outage.

What I would check:

  1. Circuit breaker state and whether it opened quickly or retries piled up.
  2. Whether a fallback provider or region is configured and was tested with the evaluation set.
  3. What the user sees: a clear message and a non-AI path (search, form, human handover).

Production consideration: Keep prompts reasonably portable, test the fallback model's quality regularly (a fallback that gives bad answers is worse than a polite refusal), and record which model served each response so post-incident review is possible.

53. A hospital discharge-summary assistant gets slower and less accurate during long sessions. Why?

Answer: Consider a hospital where doctors keep one conversation open per shift. The history keeps growing: more input tokens, higher latency, and the relevant patient context gets diluted or, worse, mixed between patients.

What I would check:

  1. How conversation IDs are chosen: per shift instead of per patient encounter is a safety problem, not just a performance one.
  2. The memory window size and whether tool results bloat history.
  3. Input token trend per request over a session.

Production consideration: Scope conversations to one patient encounter, keep a small window plus a structured summary, pull clinical facts fresh from the source system through tools instead of relying on memory, and keep a clinician review step before anything is saved to the record.

54. A GCC runs both Spring Boot and Quarkus services. Which Java AI framework do you standardise on?

Answer: There is no universal answer, so show the decision process.

What I would check:

  1. Where the AI features will actually live in the next year: mostly Spring services or both stacks.
  2. Platform requirements: observability standard, approved providers, MCP needs.
  3. Team skills and existing internal libraries.
  4. Framework maturity for the specific modules needed (some modules in both projects are newer than others).

Production consideration: A common outcome is Spring AI for Spring services and LangChain4j (via its Quarkus integration) for Quarkus services, with shared conventions above the framework: prompt storage, evaluation sets, tool authorisation rules, telemetry tags and an internal gateway for model access. Standardise the controls, not necessarily the library.

55. Quality dropped after a routine model upgrade, but all unit tests pass. What do you do?

Answer: Unit tests mock the model, so they cannot catch this; only evaluation can. Roll back the model setting first if users are affected, then investigate.

What I would check:

  1. Evaluation results on the old and new model with identical prompts and retrieval settings.
  2. Which categories regressed: format adherence, refusals, tool selection, groundedness.
  3. Whether prompts relied on quirks of the old model.
  4. Structured-output failure rate and token usage changes.

Production consideration: Treat the model ID as versioned configuration that goes through the same pipeline as code: evaluation gate, canary release, and comparison dashboards. Pin model versions where the provider allows it, instead of using a floating alias in production.

Key takeaways

  • Know the version lines: Spring AI 2.x on Spring Boot 4, Spring AI 1.x on Spring Boot 3, and what changed in tool calling and memory between them.
  • ChatClient plus advisors is the core Spring AI model; AI Services interfaces are the core LangChain4j model. The concepts map closely.
  • Tools are the main security boundary: identity from Spring Security or ToolContext, least privilege, approval for consequential actions.
  • RAG quality and RAG security both depend on metadata: tag at ingestion, filter from the authenticated user on every query.
  • LLM calls are slow I/O: use virtual threads or streaming, plus timeouts, bulkheads and fallbacks.
  • Unit tests with a mocked ChatModel prove your code; evaluation sets prove quality. You need both.
  • Prompt and completion logging is off by default for a reason; keep it that way in production unless reviewed.

Interview preparation checklist

  • Build one Spring Boot service that uses ChatClient, returns a record via entity() and calls at least two @Tool methods, one of them read-only and one behind an approval step.
  • Add RAG over PostgreSQL with pgvector, with metadata filters per department, and an ingestion job that handles updates and deletes.
  • Persist chat memory in a shared repository and explain your conversation ID strategy.
  • Write unit tests with a mocked ChatModel and an integration test with Testcontainers.
  • Add a small evaluation suite using RelevancyEvaluator or your own judge, and show results before and after a prompt change.
  • Export metrics and traces and show a dashboard with latency and token usage.
  • Expose one service method as an MCP tool, and connect to one MCP server as a client.
  • Deploy to Kubernetes with secrets from a secret store, sensible probes and streaming that works through ingress.
  • Rebuild the same feature in LangChain4j once, so you can compare the two frameworks from experience.
  • Practise scenarios Q44 to Q55 aloud, using "what I would check" and "production consideration" as your structure.

FAQ

What skills are needed for a Spring AI developer role?

Strong core Java and Spring Boot (REST, Spring Security, data access, testing), plus LLM fundamentals: prompts, structured output, tool calling, RAG with a vector store, chat memory and evaluation. Production skills such as observability, Docker, Kubernetes and secrets management separate good candidates from demo builders.

Do I need to learn Python to get a Java AI developer job?

Not for most enterprise integration roles. Spring AI and LangChain4j cover chat, tools, RAG and MCP on the JVM. Being able to read Python helps because many examples and research libraries appear there first.

Should I learn Spring AI or LangChain4j first?

If you already work with Spring Boot, start with Spring AI because it follows familiar Spring conventions. Learn LangChain4j next, especially its AI Services style, so you can discuss both in interviews.

Which Spring AI version should I practise with?

Practise with the current 2.x line on Spring Boot 4 for new projects, but understand 1.x on Spring Boot 3 too, because many company codebases are still on Boot 3. Check the official documentation for the exact current releases.

How should freshers prepare for Java AI interview questions?

Get solid at core Java, Spring Boot and SQL first, then build one small AI project with ChatClient, structured output, one tool and simple RAG. Be ready to explain every line of it and what you would add for production.

What projects should I show in a Spring Boot AI interview?

One end-to-end service with RAG, tools, tests, evaluation and observability is worth more than several chatbot demos. A domain project such as a claims triage or policy assistant, with access control on documents, gives interviewers plenty to discuss.

Are Spring AI and LangChain4j used in production?

Both have stable releases and are used for production work, but they evolve quickly and some modules are newer than others. Pin versions, read release notes before upgrading and test framework upgrades like any critical dependency.

Is Java AI development a good career path?

It is a natural extension for Java developers rather than a restart. Enterprises with large Java estates need engineers who can add AI features to existing services securely and reliably, and that combines existing backend skills with new AI skills.

How long does it take a Java developer to become interview-ready for Spring AI roles?

It depends on your current Spring Boot depth and practice time. Developers already comfortable with Spring Boot can usually cover the core Spring AI concepts and build a portfolio project within a few focused weeks; freshers should first build a solid Java and Spring foundation.

Spring AI interviews reward engineers who combine solid Java and Spring Boot skills with production AI judgement. If you want to strengthen the foundation first, the Cloudsoft Java Full Stack training covers Java, Spring Boot, REST APIs and deployment. To go further into AI, ML, cloud and security together, look at the APEX AI, ML, Cloud and Cyber Security program; and if you want to take AI systems all the way into customer environments, FDE PRO covers enterprise AI integration, deployment and evaluation, with placement support until you're placed. Classroom in Ameerpet or live online; book a free demo on +91 96660 19191.

Share𝕏infβœ‰
EnrollWhatsAppCall us