New batches starting this week Β· Limited seats

AI for Java Developers: Building LLM Features with Spring AI and LangChain4j

A practical guide for Java teams adding LLM features to Spring Boot services with Spring AI and LangChain4j, from typed structured output and tool calling to RAG, MCP, evaluation and deployment.

Adding AI to a Spring Boot service with Spring AI or LangChain4j: tools and structured output, RAG with a vector store, observability and deployment
Last updated Β· 14 min read Β· 3,156 words

AI for Java developers mostly means doing what you already do well: building secure, observable backend services, with a model call as one more dependency. Spring AI and LangChain4j give the JVM mature, production-oriented frameworks for chat calls, typed structured output, tool calling, retrieval-augmented generation (RAG) and the Model Context Protocol (MCP). For most enterprise AI application work, a Java team doesn't need to rewrite its estate in Python. It needs a handful of new patterns and the discipline to test and monitor them.

This guide is the Java counterpart to Python for AI engineers. That article covers the general habits every LLM application needs (timeouts, retries, validation, logging), so here we focus on what is specific to the JVM: the two main frameworks, how the core patterns look in Spring Boot, and how to ship them on the platforms your team already runs.

Why Java teams don't need to switch to Python

Most enterprise AI features aren't model research. They're a new capability inside an existing system: summarise this case, classify this ticket, draft a reply for a human to approve. That work lives where the data, the identity model and the transaction boundaries already are, and in banks, insurers, telecoms and many GCC engineering teams in Hyderabad and Bengaluru, that is a Spring estate.

  • The integration is the hard part. An AI feature needs customer records, entitlement checks, audit trails and core-system APIs, which your Spring services already handle.
  • JVM operations are already solved. Pipelines, images, Actuator health checks, Micrometer dashboards and runbooks exist. A separate Python service adds a second toolchain, security review and set of base images to patch.
  • Typing helps with model output. Mapping a model response into a Java record and validating it with Bean Validation is a natural fit for code that must never pass a malformed amount or account number downstream.
  • Concurrency is manageable. Model calls are I/O-bound, and virtual threads or WebFlux handle many slow outbound calls well.

A Python sidecar is reasonable when a library you need only exists there, but it shouldn't be the default.

Where Python still leads

Be honest with your team about the gaps. Python remains the home of:

  • Model training and research: PyTorch, fine-tuning toolkits, notebooks and the experiment-tracking tools data scientists use.
  • The newest libraries: new agent frameworks, evaluation tools and provider SDK features usually appear in Python (and TypeScript) first. The Java frameworks catch up, often quickly, but there is a lag.
  • Data science workflows: exploratory analysis and many document-parsing and OCR libraries.

A common split: data scientists work in Python for training and offline experiments, while product engineers build online AI features in the service's own language. Reading Python is useful for every Java developer here; rewriting production services in it usually isn't.

Spring AI vs LangChain4j

Both frameworks are open source, both have stable releases (Spring AI reached 1.0 GA in 2025 and now has a 2.x line built on Spring Boot 4; LangChain4j has been on a 1.x line since 2025, with frequent releases), and both support the major hosted model providers plus local models. Both move fast, and some modules are still marked experimental or beta.

AspectSpring AILangChain4j
HomeA Spring portfolio project, designed around Spring Boot auto-configurationAn independent Java library (not a port of Python LangChain) with Spring Boot starters and a Quarkus extension
Core programming modelChatClient, a fluent API in the style of RestClient/WebClientAI Services: you declare a Java interface and the library generates the implementation
Cross-cutting behaviourAdvisors: a chain that adds memory, RAG, guardrails or logging around each callConfigured on the AI Service builder: chat memory, content retriever, tools, moderation
Structured output.entity(MyRecord.class) maps the response to a POJO or recordThe interface method returns a POJO, record, enum or boolean
Tool calling@Tool methods passed to the prompt or set as defaults@Tool methods registered on the AI Service
RAGPortable VectorStore API (pgvector and many others), document ETL readers, RAG advisorsEmbedding stores, content retrievers, a simple "easy RAG" path and an advanced pipeline with query transformation and routing
MCPClient and server Boot starters on the official MCP Java SDK, with annotations such as @McpToolMCP client in the main project; a stdio server implementation lives in LangChain4j Community
ObservabilityBuilt-in Micrometer observations for chat, advisors, tools and vector storesModel listeners, plus a Micrometer metrics module in the Spring integration
Best fitTeams standardised on Spring Boot who want AI to feel like any other Spring dependencyTeams on Quarkus, plain Java or mixed frameworks, or who prefer the declarative interface style

You can't go badly wrong with either. In a Spring-heavy organisation, Spring AI usually wins on familiarity.

Core patterns in Java

The snippets below are short and illustrative. Class names and builder methods change between releases, so check them against the version you use.

1. A chat call and structured output into a record

The most useful first feature is usually extraction or classification: free text in, a typed object out. In Spring AI:

// Illustrative only: Spring AI (APIs evolve)
record Triage(String category, String urgency,
              boolean needsHuman, String summary) {}

@Service
class TriageService {
  private final ChatClient chat;

  TriageService(ChatClient.Builder builder) {
    this.chat = builder
        .defaultSystem("Classify the complaint. Use only "
            + "facts stated in the message.")
        .build();
  }

  Triage triage(String message) {
    return chat.prompt()
        .user(message)
        .call()
        .entity(Triage.class);
  }
}

Treat the record as untrusted input: validate it (Bean Validation, enums, range checks) and handle a mapping failure as a normal outcome. Provider-native structured output modes reduce malformed JSON but don't make the values correct. Our guide to function calling and structured outputs explains what each provider enforces.

The LangChain4j equivalent is declarative:

// Illustrative only: LangChain4j AI Service
interface Triager {
  @SystemMessage("Classify the complaint. Use only "
      + "facts stated in the message.")
  Triage triage(@UserMessage String message);
}

Triager triager = AiServices.builder(Triager.class)
    .chatModel(chatModel)
    .build();

2. Tool calling

Tools let the model ask your code to fetch data or take an action. In both frameworks you annotate a method with @Tool and a description; the framework generates the JSON schema, runs the method when the model requests it and returns the result to the model.

// Illustrative only: a read-only tool
class CaseTools {
  @Tool(description = "Status of the current "
      + "customer's open service requests")
  List<CaseStatus> openCases() {
    // customer id comes from the security context,
    // never from a model-supplied argument
    return caseClient.openCasesFor(currentCustomer());
  }
}

Two rules matter more than syntax: identity comes from your security context, never from model-filled arguments, and anything that moves money or changes a record needs an explicit approval step.

3. RAG with pgvector or another vector store

If you are new to the pattern, start with what RAG is and when to use it. In Java, the attractive option for many teams is pgvector inside a PostgreSQL instance they already run, back up and secure. Both frameworks also support dedicated stores; vector databases explained covers the trade-offs.

// Illustrative only: Spring AI RAG via an advisor
String answer = chat.prompt()
    .advisors(QuestionAnswerAdvisor
        .builder(vectorStore).build())
    .user(question)
    .call()
    .content();

Ingestion runs as a separate job that chunks, embeds and stores documents with metadata such as version and access level. Filter retrieval by that metadata so users only get chunks they may see. In LangChain4j the same flow uses an embedding store and a ContentRetriever on the AI Service builder.

4. MCP clients and servers

MCP is an open protocol, introduced by Anthropic in late 2024, for connecting AI applications to tools and data through a standard interface. Java teams meet it from two directions:

  • As a client: your Spring Boot assistant connects to existing MCP servers (an ITSM system, a code host, an internal search service) and exposes their tools to the model. Spring AI's client starter and LangChain4j's McpToolProvider both do this.
  • As a server: you expose capabilities of an existing Java service, such as "look up policy status", as MCP tools that other AI clients in the organisation can use. Spring AI's server starters let you annotate service methods and choose stdio or HTTP-based transports.

The protocol concepts are language-neutral. Our sibling tutorial on building an MCP server in Python walks through tools, resources and transports; everything there carries over, only the annotations change.

  User --> Spring Boot service --> ChatClient
                |                    |
          Spring Security       Advisors:
          (who is asking)       memory, RAG
                |                    |
          @Tool methods  <--  model requests tool
                |
       Core APIs / MCP servers / pgvector

Testing and evaluation

Split the work into two layers, because they answer different questions.

Deterministic tests (every build). Mock the chat model and test everything around it: prompt assembly, record mapping and validation, tool authorisation, fallbacks when the model times out or returns junk. Run integration tests against a real PostgreSQL with pgvector in a container (Testcontainers is the usual choice) so retrieval filters are tested on real queries. These tests are fast and shouldn't call a paid model.

Evaluation runs (on prompt, model or data changes). Keep a versioned set of realistic questions with expected outcomes, and score correctness, groundedness and tool selection. Spring AI ships an Evaluator interface with relevancy and fact-checking evaluators that use a model as a judge, which you can call from JUnit. Treat judge scores as a signal to review rather than the truth, and track results over time. The LLM evaluation guide covers datasets and metrics in more depth.

Want structured, hands-on practice with the Spring Boot, REST and microservice foundations these patterns depend on? Cloudsoft's Java Full Stack course covers that groundwork, classroom in Ameerpet or live online.

Observability with Micrometer and OpenTelemetry

Spring AI emits Micrometer observations for chat client calls, the advisor chain, model calls, tool executions and vector store operations, including token usage. Through Micrometer Tracing and an OpenTelemetry exporter, these become spans in the same distributed trace as your HTTP request and database calls. LangChain4j offers model listeners and a Micrometer metrics module for similar visibility.

What to watch, at minimum:

  • Latency per stage (retrieval, model, tools), not just end to end
  • Token usage per feature and per tenant, which is your cost signal
  • Error, timeout and fallback rates, plus output validation failures

Prompt and completion content is sensitive. Spring AI doesn't log it by default, and you should keep it that way in production unless you have masking and a retention policy that your security team has approved. For the broader design, see AI observability.

Deploying on Kubernetes

A Spring AI or LangChain4j service is still a Spring Boot or Quarkus service, so your existing Kubernetes pattern mostly applies. The AI-specific points:

  • Secrets and identity: prefer workload identity to reach cloud model endpoints (Amazon Bedrock, Azure OpenAI, Vertex AI) instead of long-lived API keys in environment variables.
  • Timeouts and probes: set explicit client timeouts, keep readiness probes independent of the model provider, and size pods for concurrent in-flight calls.
  • Egress control: restrict outbound traffic to approved model and MCP endpoints, often through an internal gateway.
  • Rollouts: put new prompts and models behind feature flags or canary releases, and treat a prompt change like a code change in the pipeline.

Kubernetes for FDE engineers goes deeper into GPU nodes, autoscaling and self-hosted models.

Example: adding an AI feature to a bank's Spring Boot service

Consider an illustrative bank GCC team in Hyderabad that owns a Spring Boot "service requests" microservice used by contact-centre agents. Agents spend a long time reading history before routing a card dispute. The team wants a suggested triage and a draft summary, with the agent always deciding.

  1. Discovery. With agents and the operations lead, they agree the feature only suggests, and define success as faster handling with no rise in misrouted cases.
  2. Structured output. A ChatClient call returns a Triage record with a category from a fixed enum, urgency and a summary. Anything that fails validation shows "no suggestion" in the UI.
  3. Read-only tools. The model can call openCases() and recentContactReasons(), both scoped to the authenticated agent's current customer through Spring Security. There are no write tools in the first release.
  4. RAG. Dispute-handling procedures are chunked and embedded into pgvector in the team's existing PostgreSQL, tagged by product and version, so the summary can cite the relevant procedure section.
  5. Data protection. Card numbers and other identifiers are masked before the prompt is built. Prompts and completions are excluded from logs; token counts and latencies are not.
  6. Evaluation. Senior agents label anonymised historical cases; every prompt or model change runs against them in CI.
  7. Deployment. The feature ships in the same Kubernetes deployment behind a flag, first to a small pilot group of agents, with a dashboard of suggestion acceptance and override rates.

Little of this is "AI code". Most of the effort is security, data handling, testing and rollout, where experienced Java engineers already add value. For the control side of such projects, see generative AI in banking.

A skills roadmap for Java developers

StageFocusProof you can show
1. FoundationsModern Java (records, virtual threads), Spring Boot, REST, Spring Security, PostgreSQL, DockerA tested, containerised CRUD service with authentication
2. LLM basicsTokens, context windows, prompting, model choice and cost; reading Python examplesA small CLI or endpoint that calls a hosted model with timeouts and error handling
3. Core patternsStructured output into records, validation, tool calling, chat memoryAn extraction service that rejects invalid output gracefully
4. RAGChunking, embeddings, pgvector, metadata filters, citationsA policy Q&A service with access-filtered retrieval
5. MCP and agentsMCP client and server, multi-step tool use, human approval stepsA Spring Boot MCP server exposing two read-only tools
6. ProductionEvaluation sets, Micrometer and OpenTelemetry, Kubernetes, CI/CD, security reviewsThe same service with an eval report, dashboards and a deployment pipeline

For a Java developer AI career, stages 1 and 6 are your head start; spend your learning time on stages 2 to 5. If you want a guided route through GenAI, RAG and agents more broadly, Cloudsoft's AI, GenAI and Agentic AI course covers them. And if what interests you is taking such systems into customer environments end to end, that is the work of a forward deployed engineer, which our FDE PRO program trains for.

FAQ

Do Java developers need to learn Python for AI?

Not to build most enterprise AI features. Spring AI and LangChain4j cover chat calls, structured output, tool calling, RAG and MCP on the JVM. Being able to read Python helps, because many examples and new libraries appear there first, but production services can stay in Java.

Should I choose Spring AI or LangChain4j?

If your services are standardised on Spring Boot, Spring AI usually feels most natural because it follows Spring conventions for configuration, auto-configuration and observability. LangChain4j suits Quarkus, plain Java or mixed estates, and teams that like declaring AI Services as interfaces.

Are Spring AI and LangChain4j production-ready?

Both have stable releases and are used for production work, but they evolve quickly and some modules are still marked experimental or beta. Pin versions, read release notes before upgrading and test them like any critical dependency.

Can I use pgvector with Spring AI and LangChain4j?

Yes. Both support PostgreSQL with the pgvector extension as a vector store, alongside many dedicated vector databases. If you already run PostgreSQL, pgvector is often the simplest start.

How do I test LLM features in a Spring Boot application?

Use two layers. Mock the model in fast unit tests to check prompt assembly, output validation, tool authorisation and fallbacks, and use Testcontainers for real database and retrieval tests. Separately, run an evaluation set of realistic questions whenever prompts, models or data change, and review the scores over time.

Can a Java service be an MCP server?

Yes. Spring AI provides MCP server starters built on the official MCP Java SDK that let you expose annotated service methods as tools over stdio or HTTP-based transports. LangChain4j focuses on the MCP client side, with a stdio server implementation in its community project.

Is AI a good career direction for Java developers?

It is a natural extension rather than a restart. Enterprises need engineers who can integrate AI features securely into existing backends, and Java developers already know the integration, security, testing and deployment side. Add RAG, tool calling, MCP and evaluation, and show them in a working project.

Strong Java, Spring Boot and microservice fundamentals are what make AI features production-ready. To build them with hands-on labs, explore Cloudsoft's Java full stack developer training, in our Ameerpet classroom or live online. Call +91 96660 19191 to book a free demo.

Share𝕏infβœ‰
EnrollWhatsAppCall us