New batches starting this week Β· Limited seats

CI/CD for AI Applications: Shipping LLM Apps Safely

In an LLM application, prompts, models, retrieval settings and tools are release artefacts. This guide walks through a CI/CD pipeline that versions, evaluates, secures and safely rolls them out.

CI/CD pipeline for AI applications: unit tests, prompt and config checks, evaluation gate, canary release, monitoring and rollback
Last updated Β· 15 min read Β· 3,205 words

CI/CD for AI applications follows the same discipline as any other software delivery: every change is versioned, tested, scanned, built once and promoted through environments. What changes is the set of things that count as a "change" and how you decide a test has passed. In an LLM application, prompts, model choices, retrieval settings and tool definitions are release artefacts just like code, and the pipeline must gate them with an evaluation suite on a fixed test set, not just unit tests. This guide walks through the pipeline stage by stage.

What is different about AI releases

A conventional web service changes when someone merges code. An LLM application can change behaviour when any of these move:

  • Prompts and system instructions. A one-line wording change can shift tone, refusals or output format everywhere.
  • Model and provider. Switching model family, model version or provider, or changing parameters such as temperature or maximum output tokens.
  • Retrieval configuration. Chunk size, overlap, embedding model, top-k, re-ranking, metadata filters and the index itself. Re-ingesting documents changes answers even if no code changed.
  • Tools and agent graphs. The model reads tool names, descriptions and schemas, so editing a description is a behavioural change; so is a new LangGraph edge or MCP server.
  • Guardrails and judge rubrics. Input/output filters and the rubrics your LLM judges score with.

The second difference is testing. LLM outputs are non-deterministic: the same input can produce different wording, occasionally a different conclusion. So the pipeline needs two kinds of tests: classic assertions for the deterministic parts (parsers, API contracts, tool wrappers, access control) and statistical evaluation for the generative parts. The evaluation methods and metrics themselves (faithfulness, answer relevance, context precision, LLM-as-judge calibration) are covered in our guide to LLM evaluation. This article is about the pipeline that runs them.

The LLMOps pipeline, end to end

Here is the reference flow most enterprise teams converge on. The names vary; the order matters.

 commit (code, prompts, configs, test set)
        |
        v
 unit tests (deterministic code)
        |
        v
 prompt / config lint + schema checks
        |
        v
 eval suite on fixed test set --fail--> block
        |
        v
 security checks (SAST, deps, secrets,
 image scan, injection test cases)
        |
        v
 build image once, sign, push
        |
        v
 deploy to staging
        |
        v
 smoke tests + canary in production
        |
        v
 full production rollout
        |
        v
 online monitoring: traces, cost, feedback
        |
        +--> new failures become test cases
StageWhat it catchesRuns when
Unit testsBroken parsing, API contracts, tool wrappers, permission logicEvery push
Prompt/config lintMissing variables, invalid YAML, unknown model IDs, bad tool schemasEvery push
Eval smoke setObvious regressions in core tasks and must-pass safety casesEvery pull request
Full eval suiteCategory-level quality drops against baselineMerge to main, release, scheduled
Security checksVulnerable dependencies, secrets, image CVEs, injection regressionsEvery pull request and build
Staging + smokeEnvironment config, IAM gaps, unreachable endpointsEvery deploy
CanaryLatency, cost and feedback under real trafficEvery production release
Online monitoringDrift, provider-side changes, new question typesContinuously

Deterministic stages run first because they are fast and free; never pay for model calls on a branch whose YAML does not parse.

Versioning prompts and configs alongside code

If a prompt lives only in a vendor dashboard or a database row that anyone can edit, you cannot answer the most important incident question: what changed? Keep every behaviour-defining artefact in the same repository as the application and review it through pull requests.

  • Prompts as files. One file per prompt in prompts/, with declared variables that lint checks against the code.
  • A single release config. One versioned file per environment pinning model ID, parameters, embedding model, retrieval and guardrail settings and prompt versions. Nothing is hardcoded.
  • Test sets in version control. Dataset, judge rubrics and baseline scores live in the repo (or a dataset store pinned from it). Changing them is a reviewed change too.
  • Index versions. Treat a vector index as a build output, recording the documents, chunking and embedding model behind it.

Prompt management in Langfuse or LangSmith helps with experimentation and linking traces to prompt versions, but production should read a pinned, reviewed version, never "latest".

Consider a GCC IT team in Hyderabad running an internal HR assistant for a global employer. A product manager edits the system prompt in a shared dashboard to make answers "friendlier". Within a day, the assistant stops adding the mandatory "check with your HR partner" line on leave-policy questions. With prompts in Git, the change would have needed a pull request, and the must-pass compliance cases would have blocked it.

Designing eval gates that people trust

An eval gate is the CI step that blocks a merge or release when quality regresses. The engineering is less about the metric library and more about agreeing what "pass" means.

Agree pass criteria with stakeholders

Thresholds should come from the people who own the risk: the business owner, compliance and the domain experts who wrote the reference answers, not a number an engineer invented in YAML. Teams typically agree three kinds of rule:

  • Must-pass cases. Safety, compliance, data-access and prompt-injection cases that must pass on every run. One failure blocks the release.
  • Relative to baseline. For scored metrics, the gate compares the candidate against the current production baseline and fails if a category drops beyond a tolerance the stakeholders signed off.
  • Human sign-off for some changes. A model or provider change, or a new tool with write access, can require a reviewer to read a sample of diffs before approval, even when the scores look fine.

Handle non-determinism in the gate

  • Use low temperature in CI where the product allows it.
  • Run known-noisy cases more than once and judge on the aggregate, not a single sample.
  • Report per category (policy questions, multi-document questions, refusals, tool-use tasks), not only an overall average, so a collapse in one area cannot hide behind gains in another.
  • Post the eval report on the pull request so reviewers see which cases changed, not just a badge.

Caching and cost control for evals in CI

Every eval case calls the model under test and often a judge too. Without controls, teams quietly stop running evals on pull requests.

  • Tier the suite. A small smoke set plus all must-pass cases on every pull request; the full test set on merge, nightly and before release.
  • Run only what changed. Use path filters: a documentation change triggers nothing, a retrieval config change triggers the RAG categories, a tool schema change triggers the agent categories.
  • Cache by content hash. Key outputs on a hash of prompt template, inputs, model ID, parameters and index version; if nothing changed for a case, reuse its output and score. Pin explicit model versions where offered, since an alias can change underneath the cache.
  • Use provider features. Prompt caching on long, shared system prompts and batch APIs for nightly full runs can reduce cost where your provider supports them.
  • Set budgets. Cap concurrency and tokens per run, and fail loudly if a run exceeds its budget instead of silently spending.

Testing tools and agents with mocks and sandboxes

Agents call tools that change real systems: create tickets, update records, send emails. CI must never touch production systems, so build two layers.

  • Mocked tools for unit and eval runs. Replace each tool with a fake that returns recorded or scripted responses. The eval then checks the agent's decisions: did it choose the right tool, pass valid arguments, stop when it should, and ask for human approval before a write action? Record tool-call traces as test fixtures.
  • Sandboxes for integration tests. Use a ServiceNow developer instance, a test Jira project, a disposable database or a containerised MCP server. Seed data, run, assert on end state, tear down.
  • Failure injection. Make mocks return timeouts, permission errors and malformed responses; agents that loop or claim success after a failed call are a common incident.

Model and provider upgrades are releases

Changing the model is the largest behavioural change you can make, yet it often arrives as a one-line config diff. Treat it with more ceremony, not less:

  1. Open a pull request that changes only the model ID and any parameters it needs.
  2. Run the full eval suite, including must-pass cases, and compare against the current baseline per category.
  3. Check cost and latency, not just quality.
  4. Have a reviewer read a sample of changed answers side by side.
  5. Roll out through canary, with the previous model still configured as the rollback target.

Also schedule the full suite when nothing in your repository changed: providers retire versions and update aliases, and documents get re-ingested. A scheduled run turns those changes into an alert rather than a customer complaint.

Canary, shadow deployments and rollback

Offline evals tell you a release is probably safe. Real traffic tells you whether it is.

  • Shadow deployment. The new version gets a copy of production requests; its answers are logged, not shown, so you compare outputs, latency and cost first. It doubles model spend for that traffic, so keep the window short, and route the shadow agent's tools to mocks or read-only modes.
  • Canary. A small share of users or one internal team gets the new version first. Watch errors, latency, cost per conversation, guardrail triggers and negative feedback by version. Argo Rollouts, a service mesh or a feature flag can split traffic.
  • Rollback. With one image and a pinned release config, rollback is pointing back to the previous versions; in GitOps, a Git revert. Keep the previous index available until the release is confirmed.

Deciding when the canary is healthy depends on the traces and dashboards described in our guide to AI observability. Running these rollouts on EKS is covered in Kubernetes for AI applications.

Pipelines like this are where many AI projects stall between a working demo and an enterprise release. If you want hands-on practice building them end to end, Cloudsoft's FDE PRO program includes GitHub Actions, Argo CD and evaluation gates inside its enterprise projects.

An illustrative GitHub Actions workflow

The YAML below is an illustrative sketch, not a drop-in file. Script names, paths and the AWS role are placeholders; adapt them to your repository and pin third-party actions to a full commit SHA in real use.

# ILLUSTRATIVE SKETCH - adapt before use
name: llm-app-ci
on:
  pull_request:
    paths: ["app/**", "prompts/**", "config/**", "evals/**"]
  push:
    branches: [main]

permissions:
  contents: read
  id-token: write   # needed for OIDC to AWS

jobs:
  test-and-lint:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pip install -r requirements.txt
      - run: pytest tests/unit
      - run: python scripts/lint_prompts.py prompts/ config/

  evals:
    needs: test-and-lint
    runs-on: ubuntu-latest
    env:
      SUITE: ${{ github.event_name == 'push' && 'full' || 'smoke' }}
    steps:
      - uses: actions/checkout@v4
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: ${{ vars.EVAL_ROLE_ARN }}
          aws-region: ap-south-1
      - uses: actions/cache@v4
        with:
          path: .eval-cache
          key: evals-${{ hashFiles('prompts/**', 'config/**') }}
      - run: pip install -r requirements.txt
      - run: python evals/run.py --suite ${{ env.SUITE }}
      - run: python evals/gate.py --rules evals/gate_rules.yaml
      - uses: actions/upload-artifact@v4
        with: { name: eval-report, path: evals/report/ }

  build:
    if: github.ref == 'refs/heads/main'
    needs: evals  # plus a security job, omitted here
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: echo "build, scan, SBOM, sign, push image"
      - run: echo "bump image tag in GitOps repo"

Note what is missing: access keys. Thresholds live in gate_rules.yaml, a reviewed file stakeholders agreed, and deployment happens through a GitOps repository that Argo CD watches.

Secrets: OIDC to the cloud, no long-lived keys

AI pipelines need credentials for model providers, vector stores, sandboxes and registries. Long-lived keys stored as CI secrets are the classic leak.

  • Federate with OIDC. GitHub Actions can exchange a short-lived OIDC token for temporary cloud credentials: an IAM role on AWS, a federated credential on a Microsoft Entra ID app registration for Azure, or Workload Identity Federation on Google Cloud. Restrict the trust policy to your repository, branch or environment.
  • Separate roles per stage. The eval role can invoke models in Amazon Bedrock or Azure OpenAI and read the eval bucket; it cannot deploy. The deploy role cannot call models.
  • Runtime secrets stay in the platform. In Kubernetes, pull provider API keys from a secrets manager at runtime instead of baking them into images or manifests.
  • Protect environments. Require reviewers on the production environment so a feature-branch workflow cannot reach production credentials.

Supply-chain checks for AI applications

AI applications pull in fast-moving SDKs, frameworks, MCP servers and sometimes model weights. Extend standard supply-chain hygiene to them.

  • Dependencies. Lock files, dependency vulnerability scanning and a review step for new packages, especially new MCP servers and agent plugins, which run with your credentials.
  • Images and actions. Minimal base images, image scanning, an SBOM per build, signed images verified at cluster admission, and GitHub Actions pinned to commit SHAs.
  • Models and data. If you self-host open-weight models, pull them from a trusted source, verify checksums and prefer safe serialisation formats. Treat ingested documents as untrusted input: they can contain prompt-injection text, so keep injection cases in the must-pass eval set.

Our guide to AI security for enterprises goes deeper on the runtime threats; the DevSecOps course covers scanning, SBOMs and policy enforcement in pipelines.

Jenkins vs GitHub Actions vs Argo CD

These are often compared as competitors, but they usually play different roles.

ToolRoleWhere it fits for AI apps
GitHub ActionsCI and workflow automation tied to the repositoryUnit tests, lint, eval gates, scans, image build; OIDC to cloud; natural choice when code lives on GitHub
JenkinsSelf-hosted, plugin-based automation serverCommon in banks and services firms with existing pipelines and on-premises agents, especially where networks restrict SaaS runners
Argo CDGitOps continuous delivery for KubernetesSyncs cluster state from Git; rollback is a revert; pairs with Argo Rollouts for canaries

A typical pattern: GitHub Actions or Jenkins runs CI and writes the new image tag into a GitOps repository; Argo CD deploys it. The eval scripts should be plain commands so they run identically in either CI system. Engineers moving from platform roles will find this familiar; our guide on the DevOps engineer to FDE transition covers what else to learn. For the tooling itself, see the Jenkins CI/CD course and the GitOps and GitHub Actions course.

An illustrative enterprise scenario

Consider an insurer deploying a claims-assistant agent that drafts responses for adjusters and can update claim notes. Prompts, release config and a test set built with claims experts live in Git, with must-pass cases such as "never state a coverage decision". Pull requests run unit tests, prompt lint and a smoke eval with mocked claim tools; merges run the full suite plus a sandbox test. When the team proposes a cheaper model, the report shows quality holding except on multi-document questions. The business owner rejects that trade-off for that category, so those questions stay on the existing model. The decision is recorded, visible and reversible: "from AI demo to enterprise outcome" in pipeline form.

Frequently asked questions

What is CI/CD for AI applications?

It is the practice of versioning, testing and releasing LLM applications through an automated pipeline, where prompts, model choices, retrieval settings and tool definitions are treated as release artefacts alongside code, and an evaluation suite gates changes in addition to unit tests.

How is an LLMOps pipeline different from a normal CI/CD pipeline?

The stages look familiar, but an LLMOps pipeline adds prompt and config linting, an evaluation suite on a fixed test set, AI-specific security cases such as prompt injection, and online monitoring of quality and cost after release. It also treats model and provider changes as releases.

How do you test non-deterministic LLM outputs in CI?

Use deterministic assertions for code and structure, and statistical evaluation for generated text: a fixed test set, low temperature where possible, repeated runs for noisy cases, per-category results and comparison against a baseline rather than exact string matching.

Who decides the pass threshold for an eval gate?

The people who own the risk: the business owner, compliance and the domain experts who wrote the reference answers. Engineers propose metrics and tolerances, but the criteria should be agreed with stakeholders and stored in a reviewed file, not invented in the pipeline.

Do I need Jenkins, GitHub Actions and Argo CD together?

No. Use one CI system, GitHub Actions or Jenkins depending on where your code and constraints are, for tests, evals and builds. Argo CD is added when you deploy to Kubernetes with GitOps, so CI updates a Git repository and Argo CD syncs the cluster.

How should CI pipelines authenticate to cloud AI services?

Use OIDC federation so the pipeline gets short-lived credentials for a narrowly scoped role, instead of storing long-lived access keys as secrets. Use separate roles for evaluation and deployment, and protect production with environment approvals.

Building pipelines that can safely ship prompts, models and agents is a core part of Forward Deployed Engineering work in customer environments. To practise it on realistic enterprise projects, take a look at the AI Forward Deployed Engineer course, available in our Ameerpet classroom or live online, with a free demo on +91 96660 19191. If you need the CI/CD and container foundations first, start with the DevOps training course.

Share𝕏infβœ‰
EnrollWhatsAppCall us