Skip to content
JackSparrow414
Go back

Automating AWS EC2 Creation, Elasticsearch and Kibana Installation, and OpenTelemetry Monitoring

Table of contents

Open Table of contents

Article body

I recently helped deploy servers. This post records and summarizes my first experience with scripted server deployment and the thinking behind it.

Why Deploy Servers with Scripts?

In production, our policy prohibits creating, bringing online, or taking offline servers by manually clicking through a cloud provider’s console. The reasons are:

  1. The process cannot be standardized. Who knows exactly what was clicked each time? A new maintainer may not know all the steps required to create a server.
  2. Without standardization, automation is impossible. Repeating the manual process for every server creates repetitive work and is highly error-prone over time.
  3. Scripts can be versioned in Git. Later, we can understand why a step was added or removed.

Choosing an EC2 Instance Type and Hardware

Before writing the script, choose the instance type and hardware for the actual workload.

Instance Type

We plan to install a standalone Elasticsearch and Kibana instance, so we choose a memory-optimized instance rather than a general-purpose, compute-optimized, or other type.

Memory

This standalone Elasticsearch instance stores non-core data. Based on that and past experience, we consider 16 GB appropriate.

CPU

According to the Elasticsearch documentation, CPU is usually not the limiting factor, so initially we believe 2 CPUs are sufficient.

Storage

We use EBS gp3 with an initial capacity of 30 GB, estimated from actual test results.

Architecture

The options are x86_64 and arm64. arm64 is somewhat cheaper, and we are gradually migrating from x86_64 to arm64, so we choose arm64.

Operating System

We standardize on Rocky Linux 8.10 across the platform. The corresponding AWS AMI is ami-06459b48b47a92d77.

Final Choice

After applying those criteria, the available instance types are r6g.large, r7g.large, and r8g.large. Since r8g.large is the newest and we are concerned about its stability, we choose the middle generation, r7g.large.

Other Configuration

Security Groups

Open Kibana port 5601 and Elasticsearch port 9200 as required to allow access from internal web servers.

Network

Use the same VPC and subnet as the other servers.

IAM Role

Configure the IAM Role as needed.

Key Pair

Use the same key pair as the other servers.

Internal Domain Name

All our servers are accessed through internal DNS names rather than IP addresses, which may change. Determine the final internal name before writing the script, for example elastic-stack-standalone.xxx.io.

Writing the Automation Script

Use AWS EC2 CLI run-instances to create the instance and AWS Route 53 CLI change-resource-record-sets to create the internal DNS record.

Tip: run-instances can create multiple instances at once; specify the number with —count.

Properties File

Put the hardware and other settings above into server.properties.

SERVER_TYPE="elastic-stack-standalone"
SERVER_INSTANCE_TYPE="r7g.large"
# arm64 rocky linux 8.9 instead of x86_64
SERVER_AMI="ami-06459b48b47a92d77"

# security group id
SG_ID="security group id"

# key pair.
KEY_PAIR_NAME=keyPairName

# networking
SUBNET_ID="subnet id"

# elastic stack standalone server does not need public IP
PUBLIC_IP=""

# private domain name
ROUTE53_FILE="change-resource-record-sets.json"
PRIVATE_DOMAIN="elastic-stack-standalone.xxx.io"
HOSTED_ZONE_ID=hostZoneId

EBS Configuration File

device-mappings.json

[
  {
    "DeviceName": "/dev/sda1",
    "Ebs": {
      "VolumeSize": 30,
      "VolumeType": "gp3",
      "DeleteOnTermination": true
    }
  }
]

EC2 Instance Creation Command

All variables except USER_DATA are read from server.properties.

aws ec2 run-instances --image-id ${SERVER_AMI} \
--key-name $KEY_NAME \
--user-data "${USER_DATA}" \
--instance-type ${SERVER_INSTANCE_TYPE} \
--block-device-mappings device-mappings.json  \
--subnet-id ${SUBNET_ID} \
--security-group-ids ${SG_ID} \
--private-ip-address $PRIVATE_IP

User Data File

User data describes the operations you want performed after AWS creates the instance. For example:

  1. Upgrade the operating system.
  2. Install software such as git and an LDAP client.
  3. Create and configure users.

Contents of user-data.txt:

install_software() {
  echo "install required software"
  yum install expect git openldap-clients sssd sssd-ldap net-tools compat-openssl10 bc -y
}
init_os() {
  # upgrade rocky linux to 8.10 from 8.9
  yum -y update
  config_security
  config_network_and_firewall
  config_system_settings_for_elastic_stack
  install_software
}

config_ldap_client() {
  echo "config ldap client"
  CONF="/git/repositories/deployment/server-setup/ldap-client"
  yes | cp -fp $CONF/etc/openldap/ldap-pro.conf /etc/openldap/ldap.conf
  yes | cp -fp $CONF/etc/sssd/sssd-pro.conf /etc/sssd/sssd.conf
  # reload sssd service
  chmod 600 /etc/sssd/sssd.conf
  systemctl restart sssd oddjobd
  systemctl enable sssd oddjobd

  # create home directory for ldap login
  authselect select sssd with-mkhomedir
  systemctl restart sshd

  #Add LDAP users to proper user groups
  for U in userList; do
    usermod -aG wheel $U
  done
}

install_elastic_stack_with_rpm() {
  rpm --import https://artifacts.elastic.co/GPG-KEY-elasticsearch

  cat <<EOF | tee /etc/yum.repos.d/elasticsearch.repo >/dev/null
[elasticsearch]
name=Elasticsearch repository for 8.x packages
baseurl=https://artifacts.elastic.co/packages/8.x/yum
gpgcheck=0
gpgkey=https://artifacts.elastic.co/GPG-KEY-elasticsearch
enabled=0
autorefresh=1
type=rpm-md
EOF
  install_elasticsearch_then_start
  install_kibana_then_start
}

install_monitor() {
   cp -rp  /root/repositories/deployment/server-setup/monitored_host /opt/
   chmod u+x /opt/monitored_host/elastic-stack/standalone/monitor.sh
   /opt/monitored_host/elastic-stack/standalone/monitor.sh
}

main() {
  init_os
  pull_git_repo
  config_ldap_client
  install_elastic_stack_with_rpm
  install_otel_monitor
}
main

See main for the overall flow. I have included only a few functions rather than every function. For example, we upgrade Rocky Linux from 8.9 to 8.10 because the latest AMI AWS offers here is 8.9, while we use 8.10. We manage server users through LDAP.

OpenTelemetry Monitoring

If you are unfamiliar with OpenTelemetry, consult its official documentation to understand what it does.

We use node_exporter for OS monitoring and elasticsearch_exporter for Elasticsearch. otel_collector collects the metrics and exposes them for the Prometheus server to scrape.

monitor.sh

#!/bin/bash
# this scirpt will install node_exporter, elasticsearch_exporter opentelemetry collector
workspace=/opt/monitored_host

architecture=$(arch)
hardware_architecture=$( [ "$architecture" = "aarch64" ] && echo "arm64" || ( [ "$architecture" = "x86_64" ] && echo "amd64" || echo "unknown-architecture" ) )

echo "The architecture is: $hardware_architecture"


install_node_exporter() {
  cd /opt
  URL=$(curl -s https://api.github.com/repos/prometheus/node_exporter/releases | grep browser_download_url | grep "linux-$hardware_architecture" | head -n 1 | cut -d '"' -f 4)
  FILE=$(echo $URL|awk -F"/" '{print $NF}')
  DIR=$(echo $URL|awk -F"/" '{print $NF}'|sed 's/\.tar\.gz//g')
  curl -LO $URL
  tar -zxf $FILE
  rm -rf /opt/node_exporter
  ln -s /opt/$DIR /opt/node_exporter
  rm -f $FILE

  # add node_exporter service
  cd $workspace
  \cp systemd_service/node_exporter.service /etc/systemd/system
  systemctl daemon-reload
  systemctl enable node_exporter.service
  systemctl start node_exporter.service
}

install_elasticsearch_exporter() {
  cd /opt
  URL=$(curl -s https://api.github.com/repos/prometheus-community/elasticsearch_exporter/releases | grep browser_download_url | grep "linux-$hardware_architecture" | head -n 1 | cut -d '"' -f 4)
  FILE=$(echo $URL|awk -F"/" '{print $NF}')
  DIR=$(echo $URL|awk -F"/" '{print $NF}'|sed 's/\.tar\.gz//g')
  curl -LO $URL
  tar -zxf $FILE
  ln -s /opt/$DIR /opt/elasticsearch_exporter
  rm -f $FILE
  # add service
  cd $workspace
  \cp systemd_service/elastic_stack_standalone_exporter.service /etc/systemd/system
  systemctl daemon-reload
  systemctl enable elastic_stack_standalone_exporter.service
  systemctl start elastic_stack_standalone_exporter.service
}

install_otelcol() {
  cd /opt
  URL=$(curl -s https://api.github.com/repos/open-telemetry/opentelemetry-collector-releases/releases|grep "browser_download_url"|grep -v "otelcol-contrib"|grep rpm|grep "linux_$hardware_architecture"|head -n 1|cut -d '"' -f 4)
  FILE=$(echo $URL|awk -F"/" '{print $NF}')
  curl -LO $URL
  rpm -iUh $FILE
  rm -f $FILE

  # add otel user
  useradd otel -s /sbin/nologin -M

  # add otel config path
  mkdir /etc/otelcol

  cd "$workspace"
  \cp elastic-stack/standalone/otelcol.yml /etc/otelcol/config.yaml
  sed -ri 's#( *host_name: ).*#\1"'$(hostname)'"#' /etc/otelcol/config.yaml
  \cp systemd_service/otelcol.service /etc/systemd/system
  systemctl daemon-reload
  systemctl enable otelcol.service
  systemctl restart otelcol.service
}
main() {
  install_n
  ode_exporter
  install_elasticsearch_exporter
  install_otelcol
}
main

otel_collector configuration:

extensions:
  health_check:

receivers:
  prometheus/os:
    config:
      scrape_configs:
        - job_name: "node_exporter"
          scrape_interval: 5s
          static_configs:
            - targets:
                - "127.0.0.1:9100"
  prometheus/elasticsearch:
    config:
      scrape_configs:
        - job_name: "elasticsearch_exporter"
          scrape_interval: 5s
          static_configs:
            - targets:
                - "127.0.0.1:9114"

exporters:
  prometheus/main:
    endpoint: "0.0.0.0:8090"
    const_labels:
      host_locale: "product"
      host_name: "replace_me"

service:
  pipelines:
    metrics/00:
      receivers: [prometheus/os, prometheus/elasticsearch]
      exporters: [prometheus/main]

Creating an Internal DNS Record

change-resource-record-sets.json

{
  "Changes": [
    {
      "Action": "UPSERT",
      "ResourceRecordSet": {
        "Name": "host.xxxx.io",
        "Type": "A",
        "TTL": 300,
        "ResourceRecords": [
          {
            "Value": "IP0"
          }
        ]
      }
    }
  ]
}

Route 53 CLI request:

# Replace values in the JSON file
ROUTE53_REQUEST=$(cat ${ROUTE53_FILE});
ROUTE53_REQUEST=${ROUTE53_REQUEST/host.${PRIVATE_DOMAIN}/${HN}.${PRIVATE_DOMAIN}}
ROUTE53_REQUEST=${ROUTE53_REQUEST/IP0/${PRIVATE_IP}}

aws route53 change-resource-record-sets --hosted-zone-id ${HOSTED_ZONE_ID} --change-batch "${ROUTE53_REQUEST}" --output json

Notifying Relevant People of Deployment Results

The deployment flow requires running only one script, after which you can work on something else. When the main script finishes, send the result to the relevant people or update the list of online servers. Integrate whichever messaging platform your company uses, as needed.

Verification

Verification should ideally also be part of the script, but I have not integrated it here. I verified manually, mainly checking:

  1. Whether the OS was upgraded.
  2. Whether the configured users can log in to the server.
  3. Whether Elasticsearch and Kibana are accessible.
  4. Whether other web servers can access Elasticsearch using its internal DNS name.
  5. Whether Prometheus can collect the metrics.

Summary

The overall approach is:

  1. Use one shared deployment script for all servers to standardize the process.
  2. Give each server its own server.properties and user-data.txt. Each deployment only needs these two files.
  3. Post-deployment verification can also be scripted and integrated into the deployment flow.

Share this post:

Continue this series

Elasticsearch and ELK in Practice

  1. Setting Up ELK and Getting Started
  2. Querying Elasticsearch
  3. Practical Elasticsearch: Common Operations, Logstash Integration, Local IP Handling, and ECS Field Mapping
  4. Generating PEM CA Certificates for ELK, Enabling HTTPS, and Connecting with the Elasticsearch Java Client
  5. Using the Elasticsearch Java API
  6. Shipping Tomcat Access Logs from EC2 to ELK with Filebeat and AWS CloudWatch Logs
  7. Shipping Tomcat access_logs from EC2 to Elasticsearch with Filebeat and AWS CloudWatch Logs, with Automated Log Management via ILM
  8. Building Elastic Stack from the Official Documentation: A Three-Node Elasticsearch Cluster, Kibana, Filebeat, Metricbeat, and Migration Without Downtime
  9. A Practical Guide to Elasticsearch in Application Development, with a Real Optimization Case
  10. Automating AWS EC2 Creation, Elasticsearch and Kibana Installation, and OpenTelemetry MonitoringYou are here
  11. Replacing Database LIKE Queries with Elasticsearch: Approaches and Implementation Details

Comments

Questions, corrections, and experiences are welcome. Sign in with GitHub to comment; both language versions share this discussion.

Comments are available on the live site only.