Skip to content
JackSparrow414
Go back

Managing and Updating Production Server Configuration in Batches with Ansible

Table of contents

Open Table of contents

Background

Our company rebuilds all production servers two or three times a year to upgrade the OS, JDK, Tomcat, and other components. Between rebuilds, other changes inevitably arise: OS security patches, or additions and changes to service configuration. Applying these changes to dozens or hundreds of servers by manually logging into each one is impractical. We therefore need to:

  1. Write the changes as scripts and commit them to Git for tracking and review.
  2. Connect to the servers over SSH and execute those scripts in batches.

Ansible does exactly this. Our Git repository ultimately contains a collection of Ansible scripts.

Set Up the Development Environment

For writing Ansible scripts locally, I recommend Ansible Development Tools with the VS Code extension. Since Python is also required, I recommend uv for isolated virtual environments. Here is a Linux example:

  1. Install uv and initialize the project.
    curl -LsSf https://astral.sh/uv/install.sh | sh
    mkdir ansible-scripts && cd ansible-scripts
    git init
    uv init
  2. Install Ansible Development Tools.
    uv tool install ansible-dev-tools
    uv run adt --version
  3. Configure VS Code and ansible-lint to detect nonstandard Ansible scripts. First install the Ansible extension.
    mkdir .vscode && cd .vscode
    vim settings.json
    Use the following settings.json:
    {
      "ansible.python.interpreterPath": "${workspaceFolder}/.venv/bin/python",
      "ansible.ansible.path": "${workspaceFolder}/.venv/bin/ansible",
      "ansible.validation.lint.path": "${workspaceFolder}/.venv/bin/ansible-lint",
      "ansible.ansibleNavigator.path": "${workspaceFolder}/.venv/bin/ansible-navigator",
      "files.associations": {
        "*.yaml": "ansible"
      },
      "[ansible]": {
        "editor.defaultFormatter": "redhat.ansible",
        "editor.formatOnSave": true,
        "editor.tabSize": 2
      },
      "yaml.validate": true,
      "yaml.format.enable": true,
      "[yaml]": {
        "editor.defaultFormatter": "redhat.vscode-yaml",
        "editor.formatOnSave": true,
        "editor.insertSpaces": true,
        "editor.tabSize": 2
      },
      "ansible.validation.lint.enabled": true,
      "ansible.validation.enabled": true
      "ansible.lightspeed.enabled": false,
      "chat.disableAIFeatures": true
    }
    Replace USER with your username. I also disabled VS Code’s AI Chat feature here.

Create an Inventory File

Among hundreds of production servers, some may host web applications, some databases, and others Redis. Each type may comprise a few to dozens of machines. First, list the classified servers in /etc/ansible/hosts. Ansible calls this inventory, with nodes organized into logical groups.

Here is an example for application servers:

[webservers]
app00
app01
app02
...
app30

The definition above can be shortened to:

[webservers]
app[00:30]

The official documentation calls this host ranges.

How do we verify that Ansible expands this notation to the expected server list? Add —list-hosts to a simple Ansible command.

ansible webservers --list-hosts

The output is:

hosts(31)
app00
app01
...
app30

For more complex inventory groups, see the official examples.

Play vs Task vs Playbook

See the official definitions.

For a simple mental model, deploying a website consists of major steps such as:

  1. Deploy the database.
  2. Deploy the web application.

Each major step is a play, and the complete workflow combining them is a playbook.

---
- name: Deploy Database
  hosts: db

  tasks:
    - name: Install MySQL
      ...

- name: Deploy Web
  hosts: web

  tasks:
    - name: Install Nginx
      ...

A play specifies which tasks to run on which hosts.

Each major step contains smaller, concrete steps:

  1. Database deployment includes installing the database, configuring users and permissions, creating tables, and loading initial SQL data required by the website.
  2. Web application deployment includes installing and configuring nginx and deploying the code.

These smaller steps are tasks within a play.

Playbook
│
├─ Play(db)
│    ├─ Task
│    ├─ Task
│    └─ Task
│
└─ Play(web)
├─ Task
├─ Task
└─ Task

A task is one operation executed on a target host.

Practical Examples

A Simple File Update

The simplest example copies a new configuration file to remote hosts, ensures the correct group ownership and permissions, and backs up the old file.

The play is:

---
- name: deploy jmx exporter configuration
  hosts: "{{ target_host | default('webservers') }}"
  become: yes

  tasks:
    - name: copy jmx_exporter.yaml to target position
      ansible.builtin.copy:
        src: ./jmx_exporter.yaml
        dest: /etc/otelcol/jmx_exporter.yaml
        owner: root
        group: root
        mode: "0644"
        backup: yes

Explanation

Pass Server Names Dynamically

The hosts value is supplied through a variable rather than hard-coded. In practice, we roll configuration changes out gradually, so the remote server names differ between runs.

Define a variable named target_host and pass it at runtime through —extra-vars.

--extra-vars "target_host=app00,app01,app02"

If I want to update ten servers each time, must I type names from app00 through app10? Since inventory supports shorthand, can this argument use shorthand too?

Ansible provides a slicing pattern: select server names by a group’s start and end indexes. The command above becomes:

--extra-vars "target_host=app[00:10]"

This updates only the eleven servers from app00 to app10. To update app15 through app25 instead:

--extra-vars "target_host=app[15:25]"

Note: slicing uses the indexes in a group’s server list, not the server names. If webservers begins with app01 rather than app00, selecting app01–05 with app[01:05] actually selects app02–06. Use app[00:04] instead.

Execute Commands as root on Remote Hosts

Only root can access the file location above, so after connecting over SSH, the current user must switch to root. Use become and become_user. become enables privilege escalation, and become_user defaults to root.

This assumes the current user belongs to the root group on the remote server.

Add -k and -K. Lowercase k prompts for the SSH password; uppercase K prompts for the sudo password on the remote host. They are usually the same, so when the become password prompt follows the SSH password prompt, simply press Enter.

-k -K

Copy Files to Remote Hosts

See the official documentation for the copy module’s options.

When starting with Ansible, you may not know which module supports a requirement. Look in the documentation’s Collection Index. Most simple requirements are covered by Ansible Builtin.

When using modules in tasks, prefer the Fully Qualified Collection Name (FQCN), as recommended in the official best practices. For a built-in module such as copy, the ansible.builtin prefix can be omitted and the module will still resolve, but including it is preferable.

For builtin modules and plugins, use the ansible.builtin collection name as a prefix, for example, ansible.builtin.copy.

Check the Script

After writing the script, use ansible-lint to check its conventions.

ansilbe-lint copy_jmx_config.yml

For other checks, see the documentation.

Run the Command

Concurrent Execution

Ansible executes plays concurrently across servers, with a default concurrency of 5. If target_host selects 00–14, execution therefore takes three rounds. To run them all together, use —forks.

--forks 15

You can also change this in the Ansible configuration.

With this background, the following command should be understandable:

Run ansible-playbook as the current user. It prompts for the remote SSH password (-k), then for the sudo password to become root (-K). Press Enter to use the same password. The script then executes the logic above concurrently on eleven servers.

ansible-playbook /home/tomcat/ansible/eng-15996/copy_jmx_config.yml --forks 11 --extra-vars "target_host=app[00:10]" -k -K

Add Configuration to a File

The next requirement is to add a datasource to Tomcat’s server.xml. Use blockinfile, which operates on blocks of text.

---
- name: Add PostgreSQL Resource to Tomcat server.xml
  hosts: "{{ target_host | default('webservers') }}"
  become: true
  become_user: tomcat

  tasks:
    - name: Insert PostgreSQL Resource into GlobalNamingResources
      ansible.builtin.blockinfile:
        path: /usr/local/tomcat/conf/server.xml
        backup: yes
        insertbefore: "</GlobalNamingResources>"
        marker: "<!-- {mark} ANSIBLE MANAGED JDBC RESOURCE -->"
        block: |
          <Resource name="jdbc/prod_postgres" username="user_dml" password="test"
                    auth="Container"
                    driverClassName="org.postgresql.Driver"
                    logAbandoned="true"
                    maxActive="20"
                    maxIdle="5"
                    minIdle="2"
                    maxWait="10000"
                    suspectTimeout="60"
                    removeAbandoned="true"
                    removeAbandonedTimeout="120"
                    abandonWhenPercentageFull="10"
                    testOnBorrow="true"
                    type="javax.sql.DataSource"
                    factory="org.apache.tomcat.jdbc.pool.DataSourceFactory"
                    url="jdbc:postgresql://test-pg-1.us-west-2.rds.amazonaws.com/prod?currentSchema=prod_app"
                    validationInterval="60000"
                    validationQuery="select 1"
                    validationQueryTimeout="3"
                 	jdbcInterceptors="org.apache.tomcat.jdbc.pool.interceptor.ConnectionState;org.apache.tomcat.jdbc.pool.interceptor.StatementFinalizer;org.apache.tomcat.jdbc.pool.interceptor.ResetAbandonedTimer"
                    defaultAutoCommit="true"/>

Explanation

Insert the block before the </GlobalNamingResources> tag in server.xml. The block contains multiple lines and uses |, the YAML multiline string syntax.

Run the Command

ansible-playbook /home/tomcat/ansible/adm-3888/add_tomcat_resource.yml --forks 10 --extra-vars "target_server=app[00:09]" -k -K

Update Configuration and Restart Services

Sometimes changing configuration is insufficient; the relevant service must be restarted through systemd to load the new configuration.

---
- name: Add local probe
  hosts: "{{ target_host | default('webservers') }}"
  become: true

  handlers:
    - name: restart custom_metrics
      ansible.builtin.systemd:
        name: custom_metrics
        state: restarted

    - name: restart otelcol
      ansible.builtin.systemd:
        name: otelcol
        state: restarted

  tasks:
    - name: Update metrics.py
      ansible.builtin.copy:
        src: ./metrics.py
        dest: /opt/common/python/custom_metrics/metrics.py
        owner: root
        group: root
        mode: "0644"
        backup: true
      notify: restart custom_metrics

    - name: Update custom_metrics.py
      ansible.builtin.copy:
        src: ./custom_metrics.py
        dest: /opt/common/python/custom_metrics/custom_metrics.py
        owner: root
        group: root
        mode: "0755"
        backup: true
      notify: restart custom_metrics

    - name: Update otelcol to add jsp metrics filter
      ansible.builtin.lineinfile:
        path: /etc/otelcol/config.yaml
        insertafter: '^\s+-\s+top_processes_memory_usage'
        line: "          - jsp_.*"
        state: present
        backup: true
      notify: restart otelcol

Explanation

Alongside the previous copy module, this example inserts a new configuration line after an existing line. For changes to individual lines, use lineinfile rather than blockinfile.

Restart Services

Use a handler: define and name it, then notify it from the task to restart the service.

Run the Command

ansible-playbook /home/tomcat/ansible/adm-5252/add_local_probe.yml --forks 7 --extra-vars "target_host=app[03:09]" -K -k

Upgrade the Linux Kernel to Fix a CVE

Once official patches are available for the recent Linux kernel vulnerability CVE-2026-31431 and related vulnerabilities, apply them promptly to production. We use Rocky 8, and the fixed kernel version is 4.18.0-553.124.

---
- name: Kernel Update and System Cleanup for CVE-2026-31431
  hosts: "{{ target_host | default('all') }}"
  become: yes
  vars:
    # CVE-2026-31431 is fixed in kernel 4.18.0-553.123 or higher
    fixed_kernel_version: "4.18.0-553.124"

  tasks:
    - name: Get current kernel version
      # Get the kernel version from ansible_facts
      set_fact:
        current_kernel: "{{ ansible_facts['ansible_kernel'] }}"

    - name: Display current kernel version
      ansible.builtin.debug:
        msg: "Current kernel version is {{ current_kernel }}"

    - name: Check if kernel is affected by CVE-2026-31431
      # Compare current version with the fixed version
      set_fact:
        is_affected: "{{ current_kernel is version(fixed_kernel_version, '<') }}"

    - name: Perform updates if the system is affected
      block:
        - name: Update all packages except temurin-17-jdk
          # dnf -y --exclude=temurin-17-jdk update
          ansible.builtin.dnf:
            name: "*"
            state: latest
            exclude: temurin-17-jdk
            update_cache: yes

        - name: Remove setroubleshoot packages
          # dnf -y remove setroubleshoot*
          ansible.builtin.dnf:
            name: "setroubleshoot*"
            state: absent

        - name: Run dnf autoremove
          # dnf -y autoremove
          ansible.builtin.dnf:
            autoremove: yes
        - name: Identify the latest kernel installed on disk
          # Query RPM to see the newest kernel version that will load after reboot
          ansible.builtin.shell: "rpm -q kernel --queryformat '%{VERSION}-%{RELEASE}.%{ARCH}\n' | tail -n 1"
          register: installed_kernel
          changed_when: false

        - name: Final status report and reboot requirement
          ansible.builtin.debug:
            msg:
              - "------------------------------------------------"
              - "PATCHING SUMMARY:"
              - "1. Currently Active Kernel: {{ current_kernel }}"
              - "2. Newly Installed Kernel: {{ installed_kernel.stdout }}"
              - "------------------------------------------------"
              - "ACTION REQUIRED:"
              - "reboot the server"
              - "------------------------------------------------"

      when: is_affected

    - name: Notify if no action is required
      ansible.builtin.debug:
        msg: "System kernel is already at or above the secure version. No patching needed."
      when: not is_affected

Explanation

Static Variables

Define fixed_kernel_version through vars. Servers already running this kernel version do not need an upgrade.

Dynamic Variables

We have not specifically tracked each server’s kernel version, even when rebuilding, so the actual versions must be discovered dynamically. Ansible provides facts for system information. This structure contains many fields; see the facts documentation. You can also run the following locally:

ansible localhost -m setup

This displays all Ansible facts for the current machine. Since Ansible supplies the information, retrieve it using the corresponding syntax.

Then use set_fact to assign the required information to current_kernel.

Compare Versions

Ansible also supplies a built-in tool for this comparison. It really is comprehensive. Use the version test; see the documentation for its syntax.

Organize Tasks Further with block

If the kernel needs updating, four subtasks must run. Use a block with a condition to group them, so the tasks in the block either all run or all skip.

Conditions

Use the basic when condition.

Capture and Print Task Output

Use register to capture task output.

Each module returns structured data containing several fields. register stores this data in the declared variable, including stdout, so later tasks can access installed_kernel.stdout directly. See Return Values for all fields.

For ordinary output, use the built-in debug module.

Run the Command

ansible-playbook /home/tomcat/ansible/adm-5240/patch_kernel.yml --extra-vars "target_server=app[03:05]" -K -k

Project Conventions

For more complex enterprise projects, Ansible provides a recommended directory layout. See Tips and Tricks for broader recommendations.

Example Project

See the source repository for a more detailed example. I built it with Kimi CodeAgent, Kimi LLM, and Superpowers. AI completed 99% of the work.

I clarified and confirmed some requirements and verified parts of the final result. Follow the documentation to start and use it; feel free to try it.

Notes

Ansible covers a great deal. When something is unclear, ask AI—or have it write the Ansible script directly. I generally let AI write these scripts now. From a learning perspective, however, understanding how we learned new tools before AI remains worthwhile and helps us use AI better in the future.


Share this post:

Previous Post
Understanding Java NIO (Part 2): I/O Multiplexing and Reactor Servers in C
Next Post
Understanding Java NIO (Part 3): I/O Multiplexing and the Reactor Pattern in Java, with Open-Source Framework Code Analysis

Comments

Questions, corrections, and experiences are welcome. Sign in with GitHub to comment; both language versions share this discussion.

Comments are available on the live site only.