Table of contents
Open Table of contents
Article body
I recently helped deploy servers. This post records and summarizes my first experience with scripted server deployment and the thinking behind it.
Why Deploy Servers with Scripts?
In production, our policy prohibits creating, bringing online, or taking offline servers by manually clicking through a cloud provider’s console. The reasons are:
- The process cannot be standardized. Who knows exactly what was clicked each time? A new maintainer may not know all the steps required to create a server.
- Without standardization, automation is impossible. Repeating the manual process for every server creates repetitive work and is highly error-prone over time.
- Scripts can be versioned in Git. Later, we can understand why a step was added or removed.
Choosing an EC2 Instance Type and Hardware
Before writing the script, choose the instance type and hardware for the actual workload.
Instance Type
We plan to install a standalone Elasticsearch and Kibana instance, so we choose a memory-optimized instance rather than a general-purpose, compute-optimized, or other type.
Memory
This standalone Elasticsearch instance stores non-core data. Based on that and past experience, we consider 16 GB appropriate.
CPU
According to the Elasticsearch documentation, CPU is usually not the limiting factor, so initially we believe 2 CPUs are sufficient.
Storage
We use EBS gp3 with an initial capacity of 30 GB, estimated from actual test results.
Architecture
The options are x86_64 and arm64. arm64 is somewhat cheaper, and we are gradually migrating from x86_64 to arm64, so we choose arm64.
Operating System
We standardize on Rocky Linux 8.10 across the platform. The corresponding AWS AMI is ami-06459b48b47a92d77.
Final Choice
After applying those criteria, the available instance types are r6g.large, r7g.large, and r8g.large. Since r8g.large is the newest and we are concerned about its stability, we choose the middle generation, r7g.large.
Other Configuration
Security Groups
Open Kibana port 5601 and Elasticsearch port 9200 as required to allow access from internal web servers.
Network
Use the same VPC and subnet as the other servers.
IAM Role
Configure the IAM Role as needed.
Key Pair
Use the same key pair as the other servers.
Internal Domain Name
All our servers are accessed through internal DNS names rather than IP addresses, which may change. Determine the final internal name before writing the script, for example elastic-stack-standalone.xxx.io.
Writing the Automation Script
Use AWS EC2 CLI run-instances to create the instance and AWS Route 53 CLI change-resource-record-sets to create the internal DNS record.
Tip: run-instances can create multiple instances at once; specify the number with —count.
Properties File
Put the hardware and other settings above into server.properties.
SERVER_TYPE="elastic-stack-standalone"
SERVER_INSTANCE_TYPE="r7g.large"
# arm64 rocky linux 8.9 instead of x86_64
SERVER_AMI="ami-06459b48b47a92d77"
# security group id
SG_ID="security group id"
# key pair.
KEY_PAIR_NAME=keyPairName
# networking
SUBNET_ID="subnet id"
# elastic stack standalone server does not need public IP
PUBLIC_IP=""
# private domain name
ROUTE53_FILE="change-resource-record-sets.json"
PRIVATE_DOMAIN="elastic-stack-standalone.xxx.io"
HOSTED_ZONE_ID=hostZoneId
EBS Configuration File
device-mappings.json
[
{
"DeviceName": "/dev/sda1",
"Ebs": {
"VolumeSize": 30,
"VolumeType": "gp3",
"DeleteOnTermination": true
}
}
]
EC2 Instance Creation Command
All variables except USER_DATA are read from server.properties.
aws ec2 run-instances --image-id ${SERVER_AMI} \
--key-name $KEY_NAME \
--user-data "${USER_DATA}" \
--instance-type ${SERVER_INSTANCE_TYPE} \
--block-device-mappings device-mappings.json \
--subnet-id ${SUBNET_ID} \
--security-group-ids ${SG_ID} \
--private-ip-address $PRIVATE_IP
User Data File
User data describes the operations you want performed after AWS creates the instance. For example:
- Upgrade the operating system.
- Install software such as git and an LDAP client.
- Create and configure users.
Contents of user-data.txt:
install_software() {
echo "install required software"
yum install expect git openldap-clients sssd sssd-ldap net-tools compat-openssl10 bc -y
}
init_os() {
# upgrade rocky linux to 8.10 from 8.9
yum -y update
config_security
config_network_and_firewall
config_system_settings_for_elastic_stack
install_software
}
config_ldap_client() {
echo "config ldap client"
CONF="/git/repositories/deployment/server-setup/ldap-client"
yes | cp -fp $CONF/etc/openldap/ldap-pro.conf /etc/openldap/ldap.conf
yes | cp -fp $CONF/etc/sssd/sssd-pro.conf /etc/sssd/sssd.conf
# reload sssd service
chmod 600 /etc/sssd/sssd.conf
systemctl restart sssd oddjobd
systemctl enable sssd oddjobd
# create home directory for ldap login
authselect select sssd with-mkhomedir
systemctl restart sshd
#Add LDAP users to proper user groups
for U in userList; do
usermod -aG wheel $U
done
}
install_elastic_stack_with_rpm() {
rpm --import https://artifacts.elastic.co/GPG-KEY-elasticsearch
cat <<EOF | tee /etc/yum.repos.d/elasticsearch.repo >/dev/null
[elasticsearch]
name=Elasticsearch repository for 8.x packages
baseurl=https://artifacts.elastic.co/packages/8.x/yum
gpgcheck=0
gpgkey=https://artifacts.elastic.co/GPG-KEY-elasticsearch
enabled=0
autorefresh=1
type=rpm-md
EOF
install_elasticsearch_then_start
install_kibana_then_start
}
install_monitor() {
cp -rp /root/repositories/deployment/server-setup/monitored_host /opt/
chmod u+x /opt/monitored_host/elastic-stack/standalone/monitor.sh
/opt/monitored_host/elastic-stack/standalone/monitor.sh
}
main() {
init_os
pull_git_repo
config_ldap_client
install_elastic_stack_with_rpm
install_otel_monitor
}
main
See main for the overall flow. I have included only a few functions rather than every function. For example, we upgrade Rocky Linux from 8.9 to 8.10 because the latest AMI AWS offers here is 8.9, while we use 8.10. We manage server users through LDAP.
OpenTelemetry Monitoring
If you are unfamiliar with OpenTelemetry, consult its official documentation to understand what it does.
We use node_exporter for OS monitoring and elasticsearch_exporter for Elasticsearch. otel_collector collects the metrics and exposes them for the Prometheus server to scrape.
monitor.sh
#!/bin/bash
# this scirpt will install node_exporter, elasticsearch_exporter opentelemetry collector
workspace=/opt/monitored_host
architecture=$(arch)
hardware_architecture=$( [ "$architecture" = "aarch64" ] && echo "arm64" || ( [ "$architecture" = "x86_64" ] && echo "amd64" || echo "unknown-architecture" ) )
echo "The architecture is: $hardware_architecture"
install_node_exporter() {
cd /opt
URL=$(curl -s https://api.github.com/repos/prometheus/node_exporter/releases | grep browser_download_url | grep "linux-$hardware_architecture" | head -n 1 | cut -d '"' -f 4)
FILE=$(echo $URL|awk -F"/" '{print $NF}')
DIR=$(echo $URL|awk -F"/" '{print $NF}'|sed 's/\.tar\.gz//g')
curl -LO $URL
tar -zxf $FILE
rm -rf /opt/node_exporter
ln -s /opt/$DIR /opt/node_exporter
rm -f $FILE
# add node_exporter service
cd $workspace
\cp systemd_service/node_exporter.service /etc/systemd/system
systemctl daemon-reload
systemctl enable node_exporter.service
systemctl start node_exporter.service
}
install_elasticsearch_exporter() {
cd /opt
URL=$(curl -s https://api.github.com/repos/prometheus-community/elasticsearch_exporter/releases | grep browser_download_url | grep "linux-$hardware_architecture" | head -n 1 | cut -d '"' -f 4)
FILE=$(echo $URL|awk -F"/" '{print $NF}')
DIR=$(echo $URL|awk -F"/" '{print $NF}'|sed 's/\.tar\.gz//g')
curl -LO $URL
tar -zxf $FILE
ln -s /opt/$DIR /opt/elasticsearch_exporter
rm -f $FILE
# add service
cd $workspace
\cp systemd_service/elastic_stack_standalone_exporter.service /etc/systemd/system
systemctl daemon-reload
systemctl enable elastic_stack_standalone_exporter.service
systemctl start elastic_stack_standalone_exporter.service
}
install_otelcol() {
cd /opt
URL=$(curl -s https://api.github.com/repos/open-telemetry/opentelemetry-collector-releases/releases|grep "browser_download_url"|grep -v "otelcol-contrib"|grep rpm|grep "linux_$hardware_architecture"|head -n 1|cut -d '"' -f 4)
FILE=$(echo $URL|awk -F"/" '{print $NF}')
curl -LO $URL
rpm -iUh $FILE
rm -f $FILE
# add otel user
useradd otel -s /sbin/nologin -M
# add otel config path
mkdir /etc/otelcol
cd "$workspace"
\cp elastic-stack/standalone/otelcol.yml /etc/otelcol/config.yaml
sed -ri 's#( *host_name: ).*#\1"'$(hostname)'"#' /etc/otelcol/config.yaml
\cp systemd_service/otelcol.service /etc/systemd/system
systemctl daemon-reload
systemctl enable otelcol.service
systemctl restart otelcol.service
}
main() {
install_n
ode_exporter
install_elasticsearch_exporter
install_otelcol
}
main
otel_collector configuration:
extensions:
health_check:
receivers:
prometheus/os:
config:
scrape_configs:
- job_name: "node_exporter"
scrape_interval: 5s
static_configs:
- targets:
- "127.0.0.1:9100"
prometheus/elasticsearch:
config:
scrape_configs:
- job_name: "elasticsearch_exporter"
scrape_interval: 5s
static_configs:
- targets:
- "127.0.0.1:9114"
exporters:
prometheus/main:
endpoint: "0.0.0.0:8090"
const_labels:
host_locale: "product"
host_name: "replace_me"
service:
pipelines:
metrics/00:
receivers: [prometheus/os, prometheus/elasticsearch]
exporters: [prometheus/main]
Creating an Internal DNS Record
change-resource-record-sets.json
{
"Changes": [
{
"Action": "UPSERT",
"ResourceRecordSet": {
"Name": "host.xxxx.io",
"Type": "A",
"TTL": 300,
"ResourceRecords": [
{
"Value": "IP0"
}
]
}
}
]
}
Route 53 CLI request:
# Replace values in the JSON file
ROUTE53_REQUEST=$(cat ${ROUTE53_FILE});
ROUTE53_REQUEST=${ROUTE53_REQUEST/host.${PRIVATE_DOMAIN}/${HN}.${PRIVATE_DOMAIN}}
ROUTE53_REQUEST=${ROUTE53_REQUEST/IP0/${PRIVATE_IP}}
aws route53 change-resource-record-sets --hosted-zone-id ${HOSTED_ZONE_ID} --change-batch "${ROUTE53_REQUEST}" --output json
Notifying Relevant People of Deployment Results
The deployment flow requires running only one script, after which you can work on something else. When the main script finishes, send the result to the relevant people or update the list of online servers. Integrate whichever messaging platform your company uses, as needed.
Verification
Verification should ideally also be part of the script, but I have not integrated it here. I verified manually, mainly checking:
- Whether the OS was upgraded.
- Whether the configured users can log in to the server.
- Whether Elasticsearch and Kibana are accessible.
- Whether other web servers can access Elasticsearch using its internal DNS name.
- Whether Prometheus can collect the metrics.
Summary
The overall approach is:
- Use one shared deployment script for all servers to standardize the process.
- Give each server its own server.properties and user-data.txt. Each deployment only needs these two files.
- Post-deployment verification can also be scripted and integrated into the deployment flow.