Memo

📈 Performance Monitoring & Tuning
📈 Performance Monitoring & Tuning
How to approach performance monitoring and tuning in Linux, and the various subsystems (and performance metrics) that need to be monitored. On a very high level, the following four subsystems need to be monitored: CPU Memory I/O Network 1. CPU Four critical performance metrics for the CPU: context switch, run queue, CPU utilization, and load average. Context Switch When the CPU switches from one process (or thread) to another, it is called a context switch. When a process switch happens, the kernel stores the current state of the CPU (of a process or thread) in memory. The kernel also retrieves the previously stored state (of a process or thread) from memory and puts it in the CPU. Context switching is essential for multitasking of the CPU. However, a higher level of context switching can cause performance issues. Run Queue The run queue indicates the total number of active processes in the current queue for the CPU. When the CPU is ready to execute a process, it picks it up from the run queue based on the priority of the process. Note that processes that are in a sleep state, or I/O wait state, are not in the run queue. A higher number of processes in the run queue can therefore cause performance issues. CPU Utilization Indicates how much of the CPU is currently being used. 100% CPU utilization means the system is fully loaded. Load Average Indicates the average CPU load over a specific time period. On Linux, load average is displayed for the last 1 minute, 5 minutes, and 15 minutes. This helps to see whether the overall load on the system is going up or down. For example, a load average of 0.75 1.70 2.10 indicates that the load is coming down (0.75 = last 1 minute, 1.70 = last 5 minutes, 2.10 = last 15 minutes). Note that this load average is calculated by combining both the total number of processes in the queue, and the total number of processes in the uninterruptible task status. 2. Network A good understanding of TCP/IP concepts is helpful when analyzing any network issue. For network interfaces, monitor the total number of packets (and bytes) received/sent through the interface, the number of packets dropped, etc. 3. I/O I/O wait is the amount of time the CPU is waiting for I/O. Consistent high I/O wait on the system indicates a problem in the disk subsystem. Monitor reads/second and writes/second. These are measured in blocks, i.e. the number of blocks read/written per second. They are also referred to as bi and bo (block in and block out). tps indicates total transactions per second, which is the sum of rtps (read transactions per second) and wtps (write transactions per second). 4. Memory RAM is the physical memory. If you have 4 GB of RAM installed, you have 4 GB of physical memory. Virtual memory = swap space available on disk + physical memory. The virtual memory contains both user space and kernel space. Using a 32-bit or a 64-bit system makes a big difference in determining how much memory a process can use: On a 32-bit system a process can only access a maximum of 4 GB of virtual memory. On a 64-bit system there is no such limitation. Unused RAM is used by the kernel as filesystem cache. Linux swaps when it needs more memory than the physical memory. When it swaps, it writes the least-used memory pages from the physical memory to the swap space on the disk. Lots of swapping can cause performance issues: the disk is much slower than the physical memory, and it takes time to swap the memory pages from RAM to disk. The subsystems are interrelated All four subsystems are interrelated. Just because you see a high reads/second, writes/second, or I/O wait, it does not mean the issue is with the I/O subsystem. It also depends on what the application is doing. In most cases, the performance issue is caused by the application running on the Linux system.
🧹 Disk Cleanup
🧹 Disk Cleanup
Find old files 1find . -type f -mtime +150 -exec ls -lrt {} \; | more 2find . -maxdepth 1 -name "*.log" -mtime +10 -exec ls -lrt {} \; 3find . -mtime +150 -exec rm -f {} \; The find loop is cheaper than a shell loop (for ...). You can make it even more efficient by batching the rm calls: 1find . -type f -print -exec rm -- "{}" + # note the "+" instead of the usual "\;" See which directories use the most space 1du -max . # list all the FS sub-directories (-x avoids filesystems other than the requested one, "." = search from where you are) 2du -sh * # show the total without listing the sub-directories (h = human readable) 3du -max . | sort -n | tail -30 # the 30 largest files/directories 4du -ks * | sort -n # size in kilobytes of all files and directories, where you are 5du -hsc * | sort -h # from smallest to largest 6ls -lrS # list files by size (in bytes) - note: ls -l does not give the true value contained in a directory 7du -ch /dir/ # size of the directories contained in /dir/ (with suffix) then the total Reduce / Truncate a file 1perl -e 'truncate "wanted_file", 100000' 2truncate -s 0 /ftpusers/ftp.upload.log File deleted but space still held by a process 1lsof +aL1 # "+L1" selects open files that have been "unlinked" (deleted but still open) 2lsof -nP | grep '(deleted)' 3find /proc/*/fd -type f -links 0 -exec ls -lrt {} \; # [SunOS] There are two solutions:
💾 Backup & Sync
💾 Backup & Sync
Rsync The classic formula 1rsync -arv --info=progress2 photo backup_photo a = archive — preserves permissions (owner, group), times, symbolic links and devices. r = recursive — copies directories and sub-directories. v = verbose — prints what is being copied. Examples 1rsync -apvz --stats --update --exclude gsast/olap_cubes --exclude gsast/param user@server-src:/export/ user@server-dest:/home/ 2rsync -av -e ssh root@192.168.1.10:/backup/DUMP/* . 3rsync -azp --stats root@oracle-src:/ec/sw/oracle/client/product/12.2.0.1/network/mesg/ ~/mesg/ 4rsync -azp /home/user/mesg/ root@oracle-dest.example.com:/ec/sw/oracle/client/product/12.2.0.1/network/mesg/ 5 6ssh root@oracle-dest.example.com "ls -lrt /ec/sw/oracle/client/product/12.2.0.1/network/mesg/" 7ssh root@oracle-dest.example.com "chown oracle:dc_dba /ec/sw/oracle/client/product/12.2.0.1/network/mesg/*" 8 9rsync -aS --delete --rsh /export/home backup-host:/export/save Propagate deletions to the backup If you delete files in the source directory, rsync does not propagate the deletion to the backup directory unless you add the --delete option.
🗜 Compression
🗜 Compression
Zip / Unzip 1zip <archive.zip> <file1> <file2> # compress files. 2zip -r <archive.zip> <directory> # compress a directory. 3unzip archive_name.zip [-d directory] # decompress an archive. Gzip / Gunzip gzip is based on the Deflate algorithm (a combination of the LZ77 and Huffman algorithms). 1gzip -l # show the size of the uncompressed file. 2gzip <file> # compress. 3gunzip <file.gz> | gzip -d <file.gz> # decompress. 4gzip -9 <my_file> # compress a file optimally. 5gzip -c <file1> <file2> > compressed_file.gz # compress several files into a single one. Bzip2 / Bunzip2 bzip2 is an alternative to gzip, more efficient but slower.
⌨️ Bash Shortcut
⌨️ Bash Shortcut
Most usefull shortcuts Ctrl + r : Reverse search. (ctrl+r to go back through the history). Ctrl + l : Clear the screen (instead of using the “clear” command). Ctrl + p : Repeat the last command. Ctrl + x + Ctrl + e : Edit the current command in an external editor (need to define export EDITOR=vim). Ctrl + shift + v : Copy / paste in Linux. Ctrl + a : Move to the beginning of the line. Ctrl + e : Move to the end of the line. Ctrl + xx : Move to the opposite end of the line. Ctrl + left : Move left one word. Ctrl + right : Move right one word.
☁️ Cloud-Init
☁️ Cloud-Init
Troubleshooting cloud-init status --wait usefull for scripting, waiting cloud-init to finish before going to next step. cloud-init status --long 1status: done 2extended_status: done 3boot_status_code: enabled-by-generator 4last_update: Thu, 01 Jan 1970 00:00:55 +0000 5detail: DataSourceNoCloud [seed=/dev/sr0] 6errors: [] 7recoverable_errors: {} sudo cloud-init analyze show 1-- Boot Record 01 -- 2The total time elapsed since completing an event is printed after the "@" character. 3The time the event takes is printed after the "+" character. 4 5Starting stage: init-local 6|`->no cache found @00.00600s +00.00000s 7|`->found local data from DataSourceNoCloud @00.01500s +00.12600s 8Finished stage: (init-local) 00.75400 seconds 9 10Starting stage: init-network 11|`->restored from cache with run check: DataSourceNoCloud [seed=/dev/sr0] @04.21100s +00.00200s 12|`->setting up datasource @04.22800s +00.00000s 13|`->reading and applying user-data @04.23400s +00.00500s 14|`->reading and applying vendor-data @04.23900s +00.00000s 15|`->reading and applying vendor-data2 @04.23900s +00.00000s 16|`->activating datasource @04.27100s +00.00100s 17|`->config-seed_random ran successfully and took 0.000 seconds @04.29500s +00.00100s 18|`->config-write_files ran successfully and took 0.001 seconds @04.29600s +00.00100s 19|`->config-growpart ran successfully and took 0.562 seconds @04.29700s +00.56200s 20|`->config-resizefs ran successfully and took 0.193 seconds @04.86000s +00.19200s 21|`->config-mounts ran successfully and took 0.001 seconds @05.05200s +00.00100s 22|`->config-set_hostname ran successfully and took 0.004 seconds @05.05300s +00.00500s 23|`->config-update_hostname ran successfully and took 0.001 seconds @05.05800s +00.00100s 24|`->config-update_etc_hosts ran successfully and took 0.005 seconds @05.05900s +00.00500s 25|`->config-users_groups ran successfully and took 0.216 seconds @05.06400s +00.21600s 26|`->config-ssh ran successfully and took 0.404 seconds @05.28100s +00.40400s 27|`->config-set_passwords ran successfully and took 0.001 seconds @05.68500s +00.00200s 28Finished stage: (init-network) 01.50000 seconds 29 30Starting stage: modules-config 31|`->config-ssh_import_id ran successfully and took 0.001 seconds @07.43300s +00.00100s 32|`->config-locale ran successfully and took 0.003 seconds @07.43400s +00.00300s 33|`->config-grub_dpkg ran successfully and took 0.352 seconds @07.43700s +00.35200s 34|`->config-apt_configure ran successfully and took 0.049 seconds @07.79000s +00.04800s 35|`->config-timezone ran successfully and took 0.007 seconds @07.83900s +00.00700s 36|`->config-runcmd ran successfully and took 0.001 seconds @07.84600s +00.00100s 37|`->config-byobu ran successfully and took 0.000 seconds @07.84700s +00.00100s 38Finished stage: (modules-config) 00.45400 seconds 39 40Starting stage: modules-final 41|`->config-package_update_upgrade_install ran successfully and took 26.632 seconds @20.56700s +26.63300s 42|`->config-write_files_deferred ran successfully and took 0.001 seconds @47.20000s +00.00200s 43|`->config-reset_rmc ran successfully and took 0.000 seconds @47.20200s +00.00100s 44|`->config-scripts_vendor ran successfully and took 0.001 seconds @47.20300s +00.00000s 45|`->config-scripts_per_once ran successfully and took 0.000 seconds @47.20300s +00.00100s 46|`->config-scripts_per_boot ran successfully and took 0.000 seconds @47.20400s +00.00000s 47|`->config-scripts_per_instance ran successfully and took 0.000 seconds @47.20400s +00.00100s 48|`->config-scripts_user ran successfully and took 0.558 seconds @47.20500s +00.55800s 49|`->config-ssh_authkey_fingerprints ran successfully and took 0.005 seconds @47.76400s +00.00500s 50|`->config-keys_to_console ran successfully and took 0.054 seconds @47.76900s +00.05500s 51|`->config-install_hotplug ran successfully and took 0.001 seconds @47.82400s +00.00100s 52|`->config-final_message ran successfully and took 0.001 seconds @47.82500s +00.00100s 53Finished stage: (modules-final) 27.29600 seconds Check the logs: sudo tail -n 50 /var/log/cloud-init-output.log
⚓ Harbor
✍️ Vim
✍️ Vim
Tutorials https://vimvalley.com/ https://vim-adventures.com/ https://www.vimgolf.com/ Plugins 1# HCL 2mkdir -p ~/.vim/pack/jvirtanen/start 3cd ~/.vim/pack/jvirtanen/start 4git clone https://github.com/jvirtanen/vim-hcl.git 5 6# Justfile 7mkdir -p ~/.vim/pack/vendor/start 8cd ~/.vim/pack/vendor/start 9git clone https://github.com/NoahTheDuke/vim-just.git Fun Facts trigger a vim tutorial vimtutor the most powerful commands: . : Repeat the last modification. * : Where the cursor is located, keeps the word in memory and goes to the next occurrence. .* : together, repeat an action on the next word.
🌅 UV
🌅 UV
Install 1# curl method 2curl -LsSf https://astral.sh/uv/install.sh | sh 3 4# Pip method 5pip install uv Quick example 1pyenv install 3.12 2pyenv local 3.12 3python -m venv .venv 4source .venv/bin/activate 5pip install pandas 6python 7 8# equivalent in uv 9uv run --python 3.12 --with pandas python Usefull 1uv python list --only-installed 2uv python install 3.12 3uv venv /path/to/environment --python 3.12 4uv pip install django 5uv pip compile requirements.in -o requirements.txt 6 7uv init myproject 8uv sync 9uv run manage.py runserver Run as script Put before the import statements: 1#!/usr/bin/env -S uv run --script 2# /// script 3# requires-python = ">=3.12" 4# dependencies = [ 5# "ffmpeg-normalize", 6# ] 7# /// Then can be run with uv run sync-flickr-dates.py. uv will create a Python 3.12 venv for us. For me this is in ~/.cache/uv (which you can find via uv cache dir).
🌐 Network troubleshooting
🌐 Network troubleshooting
Troubleshoot DNS vi dns.yml 1apiVersion: v1 2kind: Pod 3metadata: 4 name: dnsutils 5 namespace: default 6spec: 7 containers: 8 - name: dnsutils 9 image: registry.k8s.io/e2e-test-images/jessie-dnsutils:1.3 10 command: 11 - sleep 12 - "infinity" 13 imagePullPolicy: IfNotPresent 14 restartPolicy: Always deploy dnsutils 1k apply -f dns.yml 2pod/dnsutils created 3 4kubectl get pods dnsutils 5NAME READY STATUS RESTARTS AGE 6dnsutils 1/1 Running 0 36s Troubleshoot with dnsutils 1kubectl exec -i -t dnsutils -- nslookup kubernetes.default 2;; connection timed out; no servers could be reached 3command terminated with exit code 1 4 5kubectl exec -ti dnsutils -- cat /etc/resolv.conf 6search default.svc.cluster.local svc.cluster.local cluster.local example.local 7nameserver 10.43.0.10 8options ndots:5 9 10kubectl get endpoints kube-dns --namespace=kube-system 11NAME ENDPOINTS AGE 12kube-dns 10.42.0.6:53,10.42.0.6:53,10.42.0.6:9153 5d1h 13 14kubectl get svc kube-dns --namespace=kube-system 15NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE 16kube-dns ClusterIP 10.43.0.10 <none> 53/UDP,53/TCP,9153/TCP 5d1h CURL 1cat << EOF > curl.yml 2apiVersion: v1 3kind: Pod 4metadata: 5 name: curl 6 namespace: default 7spec: 8 containers: 9 - name: curl 10 image: curlimages/curl 11 command: 12 - sleep 13 - "infinity" 14 imagePullPolicy: IfNotPresent 15 restartPolicy: Always 16EOF 17 18k apply -f curl.yml 19 20#Test du DNS 21kubectl exec -i -t curl -- curl -v telnet://10.43.0.10:53 22kubectl exec -i -t curl -- curl -v telnet://kube-dns.kube-system.svc.cluster.local:53 23kubectl exec -i -t curl -- nslookup kube-dns.kube-system.svc.cluster.local 24 25curl -k -I --resolve subdomain.domain.com:52.165.230.62 https:/subdomain.domain.com/
🎡 Helm
🎡 Helm
Administration See what is currently installed 1helm list -A 2NAME NAMESPACE REVISION UPDATED STATUS CHART APP VERSION 3nesux3 default 1 2022-08-12 20:01:16.0982324 +0200 CEST deployed nexus3-1.0.6 3.37.3 Install/Uninstall 1helm status nesux3 2helm uninstall nesux3 3helm install nexus3 <chart> # chart URL or path 4helm history nexus3 5 6# work even if already installed 7helm upgrade --install ingress-nginx ${DIR}/helm/ingress-nginx \ 8 --namespace=ingress-nginx \ 9 --create-namespace \ 10 -f ${DIR}/helm/ingress-values.yml 11 12#Make helm unsee an apps (it does not delete the apps) 13kubectl delete secret -l owner=helm,name=argo-cd Handle Helm Repo and Charts 1#Handle repo 2helm repo list 3helm repo add gitlab https://charts.gitlab.io/ 4helm repo update 5 6#Pretty usefull to configure 7helm show values elastic/eck-operator 8helm show values grafana/grafana --version 8.5.1 9 10#See different version available 11helm search repo hashicorp/vault 12helm search repo hashicorp/vault -l 13 14# download a chart 15helm fetch ingress/ingress-nginx --untar
🎫 Certificates Authority
🎫 Certificates Authority
Trust a CA on Linux host 1# [RHEL] RootCA from DC need to be installed on host: 2cp my-domain-issuing.crt /etc/pki/ca-trust/source/anchors/my_domain_issuing.crt 3cp my-domain-rootca.crt /etc/pki/ca-trust/source/anchors/my_domain_rootca.crt 4update-ca-trust extract 5 6# [Ubuntu] 7sudo apt-get install -y ca-certificates 8sudo cp local-ca.crt /usr/local/share/ca-certificates 9sudo update-ca-certificates