ScreenshotNeo

BlogEngineering

Jenkins Best Practices for Reliable CI/CD

A practical Jenkins reliability guide to pipeline-as-code, agent capacity, credentials, backups, durability settings, upgrades, and recovery.

By the ScreenshotNeo team4 October 202611 min read

Reliable Jenkins CI/CD comes from making pipeline changes reviewable, keeping build execution off the controller, limiting credential exposure, practicing recovery, and testing upgrades before production. Start with these five controls; then tune agent capacity and Pipeline durability to match your workloads and recovery needs.

The Jenkins handbook groups its guidance around automating job definitions, managing jobs, and managing the controller. The recommendations below follow that operational view. Exact settings depend on your Jenkins and plugin versions, infrastructure, and the cost of rerunning or recovering a build.

1. Store pipeline definitions in source control

Put a Jenkinsfile in the repository with the application code. Jenkins recommends this approach because the Pipeline can be reviewed, iterated on, audited, and shared as a single definition. A change to build or deployment behavior then follows the same review and history as a code change. See Using a Jenkinsfile.

Declarative Pipeline is a useful default when its structured syntax covers the workflow. A Declarative Pipeline requires an agent; it organizes work into stages and steps. Here is a complete example for a repository with a shell-based test command and a deploy script:

pipeline {
    agent { label 'linux' }

    options {
        timestamps()
        disableConcurrentBuilds()
        timeout(time: 30, unit: 'MINUTES')
    }

    stages {
        stage('Checkout') {
            steps {
                checkout scm
            }
        }
        stage('Test') {
            steps {
                sh './scripts/test.sh'
            }
        }
        stage('Package') {
            steps {
                sh './scripts/package.sh'
            }
        }
        stage('Deploy') {
            when {
                branch 'main'
            }
            steps {
                withCredentials([string(credentialsId: 'deploy-token', variable: 'DEPLOY_TOKEN')]) {
                    sh 'set +x; ./scripts/deploy.sh'
                }
            }
        }
    }

    post {
        always {
            archiveArtifacts artifacts: 'dist/**', allowEmptyArchive: true
        }
        failure {
            echo 'Pipeline failed; inspect the stage logs and notify the owning team.'
        }
    }
}

Replace the label, commands, credential ID, artifact path, and branch condition with values that exist in your installation. This example assumes the scripts are checked into the repository and that the Jenkins credential is scoped to the job or folder that needs it. Do not copy a deploy stage into a job that runs untrusted pull request code.

Keep job creation and branch discovery consistent

For repositories with multiple branches, use multibranch Pipeline or organization folders where they fit your source control setup. These approaches let Jenkins discover branches and use the repository’s Pipeline definition. Keep job names simple and consistent, and make failures visible to the people who can act on them. Jenkins’ Best Practices also recommends avoiding the legacy Maven project job type in favor of Pipeline-based approaches.

2. Run builds on agents and size capacity deliberately

The controller coordinates jobs, manages configuration, and schedules work. Agents execute the build steps. Jenkins explicitly recommends agents instead of running builds on the controller; its agent guidance recommends setting the controller’s executor count to zero. See Using Jenkins agents and Controller Isolation.

Agents provide capacity and help isolate the controller from resource-heavy or untrusted build scripts. They do not automatically make a system secure: an agent with broad permissions, access to controller files, or access to powerful credentials can still create risk. Treat agent trust boundaries as part of the design.

Choose labels and executors to fit workloads

  • Use labels to direct jobs to compatible environments, such as linux, windows, or a specific architecture or toolchain.
  • Start with a conservative executor count. Jenkins’ node guidance describes one executor per node as the safest configuration; more may work for small tasks, but monitor CPU, memory, disk and I/O when increasing concurrency.
  • Separate workloads when they need different operating systems, hardware, trust levels, or resource limits.
  • Use ephemeral or dynamically provisioned agents when that matches your infrastructure and operating model. Ensure required tools and dependencies are available reproducibly rather than relying on undocumented state left on a worker.
  • Keep enough agent capacity for the normal arrival rate and build duration. When the queue grows, check whether jobs are waiting for a matching label, an executor, a lock, or an external dependency before simply adding executors.

Jenkins describes the controller as the orchestration hub and agents as build execution machines in its scale architecture guidance. Multiple controllers can isolate teams or workloads, but they also multiply the work of securing, upgrading, and backing up Jenkins. Choose that topology based on project criticality and the organization’s capacity to operate it; it is not a universal prerequisite.

3. Protect credentials and untrusted Pipeline code

Keep Jenkins security enabled, grant permissions narrowly, and define credentials at the lowest suitable scope. A credential available to a folder or job should not be made global without a reason. Review who can configure jobs and who can change the repository files that Jenkins executes.

Bind secrets only around the steps that need them. Avoid printing secret values, passing them as command-line arguments when that could expose them in process listings, or interpolating secrets in Groovy strings. Jenkins’ Jenkinsfile credential examples use single-quoted shell scripts so the shell expands environment variables instead of Groovy interpolating the secret into the command string.

withCredentials([string(credentialsId: 'api-token', variable: 'API_TOKEN')]) {
    sh '''
        set +x
        ./scripts/call-service.sh
    '''
}

The script can read API_TOKEN from its environment. Ensure it does not echo the value or include it in error output. Masking in console logs can reduce accidental disclosure, but it is not a security boundary: Pipeline code that can access a credential can intentionally transmit or otherwise expose it. Do not make trusted credentials available to untrusted Pipeline jobs, such as jobs that execute pull request code you have not approved. Jenkins explains these limits in its credential guidance and the credentials masking limitations article.

Separate trust levels

  • Decide which branches and contributors may run deployment or signing steps.
  • Keep privileged credentials out of jobs that execute unreviewed changes.
  • Restrict agent permissions and network access to what the job requires.
  • Review plugin and job configuration permissions as part of the same access-control model as repository permissions.

4. Make backups recoverable, not merely present

A backup is useful only if it contains the needed state and can be restored. Define what to save, how often to save it, where copies live, and the recovery objective the team expects. Jenkins’ backup and restore documentation recommends validating backups, including by restoring one to a temporary location.

  1. Inventory the controller data and configuration your recovery plan needs, including JENKINS_HOME, job definitions, plugin and tool state where required, and build records your team must retain.
  2. Choose a backup method and schedule that fit the amount of change and the acceptable data-loss window. Filesystem snapshots are one option where the storage platform supports them.
  3. Protect backup copies from the same failure or access path as the controller, according to your organization’s security requirements.
  4. Keep the Jenkins controller key separate from routine backups and store it securely. Recovery procedures need to restore it separately so encrypted data remains usable.
  5. Periodically restore a copy in an isolated temporary location. Check that Jenkins starts, jobs and required configuration are present, and the restore procedure is documented and repeatable.

A basic validation launch, adapted to your environment and Jenkins distribution, can use a separate home directory and an unused HTTP port:

export JENKINS_HOME=/mnt/backup-test
java -jar jenkins.war --httpPort=9999

Do not point a test restore at the production home or port. Confirm the right Jenkins and plugin versions for the recovery plan, and avoid relying on a backup process that has never been exercised end to end.

5. Set Pipeline durability according to recovery needs

Pipeline durability trades disk I/O against recovery of in-flight Pipeline state after an abrupt controller stop. Jenkins says Pipelines can survive planned and unplanned controller restarts, but the result depends on durability settings and how the shutdown occurs. See Scaling Pipelines.

Setting Trade-off Typical fit
Maximum survivability (MAX_SURVIVABILITY) Writes Pipeline state more consistently and is the slowest option. Critical pipelines where retaining in-flight execution state or an execution record matters.
Performance-optimized (PERFORMANCE_OPTIMIZED) Reduces disk I/O; an abrupt shutdown can leave in-flight pipelines unable to resume or display full state. Rerunnable build and test work where the team accepts restarting an interrupted run.
Less durable (SURVIVABLE_NONATOMIC) Writes state at each step but avoids atomic writes; faster in some environments with a small additional risk. Workloads balancing I/O cost with a need to record steps.

Durability can be set globally, per Pipeline job, or for multibranch projects; more specific settings can override the global default. Changes apply to subsequent applicable runs, not an already-running Pipeline. Verify the available options and required Pipeline plugin versions against the Jenkins version you operate. A reasonable policy is to keep the more durable mode for critical deployments or audit-sensitive runs and consider faster settings for work that is safe to rerun.

6. Test Jenkins core and plugin upgrades before rollout

Core and plugin upgrades can change compatibility and behavior. Jenkins warns that an upgrade may impair another plugin or crash a controller. Maintain a test deployment that represents the production environment closely enough to exercise important jobs before rolling changes out. Consult the official Jenkins upgrade guides and plugin documentation for the versions you plan to use.

  1. Record the current Jenkins core and plugin versions, plus relevant configuration and agent images.
  2. Read upgrade notes and check plugin compatibility for the target core version.
  3. Restore or reproduce a representative controller in a test environment.
  4. Run important pipelines, including checkout, tests, artifact handling, notifications, and deployment paths in a safe environment.
  5. Plan a rollout and recovery path. Keep a known-good backup and document which version set it represents.

Do not assume that a successful controller startup proves the upgrade is safe: exercise the jobs and integrations the team depends on. Conversely, a test environment only reduces risk if it is maintained closely enough to reveal relevant differences.

7. Troubleshoot common reliability failures

Symptom Likely cause What to check or change
Jobs queue but do not start No online agent matches the label, all matching executors are busy, or a job is waiting on a lock or resource. Inspect the queue reason, agent status, labels, executor availability, and any lock or throttling configuration. Add capacity only after identifying the bottleneck.
Controller becomes slow during builds Builds are using controller executors or controller resources are contended by job activity. Move execution to agents and set built-in node executors to zero. Review controller CPU, memory, disk latency, and workload.
Agent repeatedly goes offline Connectivity, Java/runtime compatibility, resource exhaustion, disk thresholds, or agent process failures. Check the agent log and controller node status; verify network reachability and supported runtime, and monitor disk, temporary space, memory, and CPU.
Pipeline cannot resume after a restart The controller stopped abruptly while a less durable mode had not persisted all in-flight state. Check the Pipeline durability setting and shutdown circumstances. Rerun safe work; use a more durable setting for critical workflows where state recovery is required.
Secret appears in logs or command output Pipeline code printed it, Groovy interpolated it, shell tracing was enabled, or a transformed value escaped masking. Revoke or rotate the exposed credential, inspect access, stop logging it, disable tracing around the command, use safe credential binding, and review the Jenkins credential guidance.
Backup exists but restoration fails Required data or the controller key is missing, versions do not match the recovery plan, or the backup is incomplete. Restore a test copy, validate the key handling and required files, and update the recovery procedure before relying on the next backup.
Plugin upgrade breaks jobs or startup Incompatible plugin/core versions or behavior changes. Use the test deployment to identify the failing dependency, consult upgrade notes, and follow the documented recovery plan.
Builds fail only on busy agents Concurrency exceeds available CPU, memory, disk I/O, or another shared resource. Reduce executors or isolate the workload, then monitor resource use under representative load before increasing concurrency again.

8. Improve reliability without hiding failures

  • Use explicit timeouts for jobs and external operations that could otherwise wait indefinitely.
  • Make retry behavior selective. Retry transient network or service failures when safe; do not turn deterministic test or compilation errors into silent repeated work.
  • Keep artifacts and logs useful for diagnosis, while applying retention policies appropriate to storage limits and audit needs.
  • Notify the team responsible for a failure, with enough context to find the run and stage. Avoid sending every routine status to people who cannot act on it.
  • Prefer reproducible build inputs and explicit tool versions so agent replacement does not depend on undocumented machine state.
  • Monitor queue time, agent availability, controller health, disk capacity, and backup validation outcomes. Use trends to identify capacity or reliability problems before they become incidents.

These are operating practices, not guarantees that every failure can be prevented. Define which failures are acceptable to rerun, what state must be preserved, and who owns restoring the service.

Or skip the browser setup

Jenkins teams often need screenshots of build dashboards, job pages, or deployment checks for release records and visual review. ScreenshotNeo can capture a page with one GET request. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://jenkins.io -o shot.webp

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, timeouts, and failed loads are not billed, and cache hits cost nothing; response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is available on every plan. Learn more at ScreenshotNeo, then sign up for 1,000 free screenshots a month with no card.

Jenkins reliability checklist

  • Pipeline definitions live in source control and receive review.
  • Build execution runs on appropriately labeled agents; controller executors are set to zero.
  • Executor counts reflect measured workload capacity.
  • Credentials are scoped narrowly, and untrusted jobs cannot use trusted secrets.
  • Backups include the required state, protect the controller key separately, and have been restored in a test location.
  • Pipeline durability settings match the recovery needs of each workload.
  • Core and plugin upgrades are exercised in a representative test deployment.
  • Queue, agent, controller, storage, and backup health have clear owners.

FAQ

Should every team use multiple Jenkins controllers?

No. Separate controllers can isolate workloads, but each one adds operational work. Choose based on criticality, isolation requirements, and the team’s ability to secure, update, and recover each controller.

Does agent execution make Jenkins jobs safe by itself?

No. It keeps build work off the controller, but agents still need suitable permissions, network boundaries, and credential access controls. Treat build scripts and contributors according to their trust level.

Can I use performance-optimized durability for deployments?

Only if the team accepts the recovery trade-off. For deployments or workflows where a complete execution record or in-flight state is critical, consider a more durable setting and verify its behavior with your installed versions.

How often should I test a backup restore?

Set a cadence based on how often controller state changes and the recovery objective. The key requirement is to validate periodically and after material changes to the backup or restore process.