πŸ“ Scripting
πŸ“ Scripting
Inside a Shell script One line command: 1# Set the SID 2ORAENV_ASK=NO 3export ORACLE_SID=orcl 4. oraenv 5 6# Trigger oneline command 7echo -e "select inst_id, instance_name, host_name, database_status from gv\$instance;" | sqlplus -S / as sysdba In bash script: 1su - oracle -c ' 2export SQLPLUS="sqlplus -S / as sysdba" 3export ORAENV_ASK=NO; 4export ORACLE_SID='${SID}'; 5. oraenv | grep -v "remains"; 6 7${SQLPLUS} <<EOF2 8set lines 200 pages 2000; 9select inst_id, instance_name, host_name, database_status from gv\$instance; 10exit; 11EOF2 12 13unset ORAENV_ASK; 14' Inside SQL Prompt 1-- with an absolute path 2@C:\Users\Matthieu\test.sql 3 4-- or trigger from director on which sqlplus was launched 5@test.sql 6 7-- START syntax possible as well 8START test.sql Variables usages 1-- User variable (if not define, oracle will prompt) 2SELECT * FROM &my_table; 3 4-- Prompt user to set a variable 5ACCEPT my_table PROMPT "Which table would you like to interrogate ? " 6SELECT * FROM $my_table; Some Examples Example of Shell script to launch sqlplus command: 1export ORACLE_SID=SQM2DWH3 2 3echo "connect ODS/ODS 4BEGIN 5ODS.PURGE_ODS.PURGE_LOG(); 6ODS.PURGE_ODS.PURGE_DATA(); 7END; 8/" | sqlplus /nolog 9 10echo "connect DSA/DSA 11BEGIN 12DSA.PURGE_DSA.PURGE_LOG(); 13DSA.PURGE_DSA.PURGE_DATA(); 14END; 15/" | sqlplus /nolog Example of script to check tablespaces.sh 1#!/bin/ksh 2 3sqlplus -s system/manager <<! 4SET HEADING off; 5SET PAGESIZE 0; 6SET TERMOUT OFF; 7SET FEEDBACK OFF; 8SELECT df.tablespace_name||','|| 9 df.bytes / (1024 * 1024)||','|| 10 SUM(fs.bytes) / (1024 * 1024)||','|| 11 Nvl(Round(SUM(fs.bytes) * 100 / df.bytes),1)||','|| 12 Round((df.bytes - SUM(fs.bytes)) * 100 / df.bytes) 13 FROM dba_free_space fs, 14 (SELECT tablespace_name,SUM(bytes) bytes FROM dba_data_files GROUP BY tablespace_name) df 15 WHERE fs.tablespace_name (+) = df.tablespace_name 16 GROUP BY df.tablespace_name,df.bytes 17 ORDER BY 1 ASC; 18quit 19! 20 21exit 0 1#!/bin/ksh 2 3sqlplus -s system/manager <<! 4 5set pagesize 60 linesize 132 verify off 6break on file_id skip 1 7 8column file_id heading "File|Id" 9column tablespace_name for a15 10column object for a15 11column owner for a15 12column MBytes for 999,999 13 14select tablespace_name, 15'free space' owner, /*"owner" of free space */ 16' ' object, /*blank object name */ 17file_id, /*file id for the extent header*/ 18block_id, /*block id for the extent header*/ 19CEIL(blocks*4/1024) MBytes /*length of the extent, in Mega Bytes*/ 20from dba_free_space 21where tablespace_name like '%TEMP%' 22union 23select tablespace_name, 24substr(owner, 1, 20), /*owner name (first 20 chars)*/ 25substr(segment_name, 1, 32), /*segment name */ 26file_id, /*file id for extent header */ 27block_id, /*block id for extent header */ 28CEIL(blocks*4/1024) MBytes /*length of the extent, in Mega Bytes*/ 29from dba_extents 30where tablespace_name like '%TEMP%' 31order by 1, 4, 5 32/ 33 34quit 35! 36 37exit 0 SPOOL to write on system from sqlplus: 1SQL> SET TRIMSPOOL on 2SQL> SET LINESIZE 1000 3SQL> SPOOL /root/output.txt 4SQL> select RULEID as RuleID, RULENAME as ruleName,to_char(DBMS_LOB.SUBSTR(EPLRULESTATEMENT,4000,1() as ruleStmt from gep_rules; 5SQL> SPOOL OFF from script.sql: 1SET TRIMSPOOL on 2SET LINESIZE 10000 3SPOOL resultat.txt 4ACCEPT var PROMPT "Which table do you want to get ? " 5SELECT * FROM &var; 6SPOOL OFF Generate DATA Duplicate table to fill up tablespace or generate fake data: 1SQL> Create table emp as select * from employees; 2SQL> UPDATE emp SET LAST_NAME='ABC'; 3SQL> commit;
πŸ“¦ Export / Import (Data Pump)
πŸ“¦ Export / Import (Data Pump)
Export EXP (legacy): the old export utility β€” produces a binary dump (superseded by Data Pump). EXPDP (Data Pump): produces binary dump files, used with DIRECTORY objects. 1exp user/password@host FULL=Y # full legacy (binary) export. 2expdp user/password@host FULL=Y DIRECTORY=DUMP DUMPFILE=full.dmp 1# full DB, excluding statistics 2nohup expdp 'system/<password>'@orcl FULL=Y DIRECTORY=DUMP \ 3 DUMPFILE=expdp_orcl_full_$(date +%Y-%m-%d).dmp \ 4 LOGFILE=expdp_orcl_$(date +%Y-%m-%d).log EXCLUDE=statistics 5 6# one schema 7nohup expdp system/<password>@orcl SCHEMAS=my_schema DIRECTORY=DUMP \ 8 DUMPFILE=my_schema_$(date +%Y-%m-%d)_%U.dmp \ 9 LOGFILE=my_schema.log EXCLUDE=statistics Directories & rights 1SET LINES 200 PAGES 2000 2SELECT * FROM dba_directories; -- the paths defined for Oracle. 1CREATE DIRECTORY my_dir AS '/backup/dump'; 2GRANT READ, WRITE ON DIRECTORY my_dir TO my_user; 3DROP DIRECTORY my_dir; Then use it in an export: ... DIRECTORY=my_dir DUMPFILE=my_export.dmp.
πŸ”§ Installation
πŸ”§ Installation
Sources & Docs Oracle-Base: DB 19c RAC installation on Oracle Linux 8 (VirtualBox) Standards Keep every Oracle installation as uniform as possible (easier automation). The points below are all required. Example migration: previous install RAC ONE NODE SE (Grid 19.0.0 + udev/ASM, DB 12.2.0.1) β†’ new RAC Active/Active EE (Grid 19.3 + AFD/ASM, DB 19.10). Users & groups 1grep oracle /etc/passwd # oracle:x:1521:1521:Oracle User For Database Binaries:/home/oracle:/bin/bash 2grep oinstall /etc/group # oinstall:x:1521:oracle 3grep dba /etc/group # dba:x:1522:oracle Filesystems & diskgroups /u01 β€” a dedicated 100G filesystem (binaries ~25G + full install ~15G). /tmp β€” minimum 4G. DATA β€” ~60G raw devices / disks. FRA β€” minimum 4 disks Γ— 20G (or 40G), raw devices / disks. VOT β€” minimum one 5G disk for the voting disk. Network Single instance: minimum two interfaces (public + backup/NFS). RAC: three interfaces: 1DEVICE TYPE CONNECTION 2ens192 ethernet Admin 3ens224 ethernet Interconnect 4ens256 ethernet Backup One network interface for backups is required to mount an NFS share.
πŸ—„οΈ Tablespace
πŸ—„οΈ Tablespace
Concepts A tablespace (TBS) is a logical group of storage for data; each tablespace is made of one or more datafiles (.dbf), created on a disk (TBS = 1.dbf + 2.dbf + …). One datafile belongs to exactly one tablespace; a tablespace can have many datafiles. To grow a database you grow the datafiles of the required tablespace (which needs free space on the filesystem or ASM disk). View tablespaces & datafiles 1SELECT * FROM dba_tablespaces; -- the tablespaces. 2SELECT * FROM dba_data_files; -- the datafiles. 3SELECT * FROM dba_temp_files; -- the temporary files. 1SELECT tablespace_name FROM dba_tablespaces; Size & max size of a tablespace (interactive):
πŸ› οΈ Administrations
πŸ› οΈ Administrations
Identify the instance 1echo $ORACLE_SID # the Oracle instance used before a SQL connection. 2echo $ORACLE_HOME # the Oracle home directory. Load the instance environment: 1. oraenv /etc/oratab: 1+ASM:/u01/oracle/base/product/12.2.0/grid:N 2ORCL:/u01/oracle/base/product/12.2.0/dbhome_1:N All running instances: 1ps -ef | grep pmon 2oracle 2201 1 0 12:02 ? 00:00:00 ora_pmon_ORCL 3oracle 30513 1 0 Feb12 ? 00:01:11 asm_pmon_+ASM 1ps -ef | grep ora_pmon | grep -v grep | awk '{print $NF}' | cut -d"_" -f3 With a Clusterware layer Check whether an Oracle Clusterware layer is present: 1ps -ef | grep d.bin 2# /u01/oracle/base/product/12.2.0/grid/bin/[ohasd|oraagent|evmd|ocssd].bin ... 1srvctl config database Listeners and processes:
🧩 Clusterware
🧩 Clusterware
Grid The grid is the component responsable for Clustering in oracle. Grid (couche clusterware) -> ASM -> Disk Group - Oracle Restart = Single instance = 1 Grid (with or without ASM) - Oracle RAC OneNode = 2 instances Oracle in Actif/Passif with shared storage - Oracle RAC (Actif/Actif) SCAN 1# As oracle user: 2srvctl config scan 3 4SCAN name: host-env-datad1-scan.domain, Network: 1 5Subnet IPv4: 192.168.228.0/255.255.255.0/ens192, static 6Subnet IPv6: 7SCAN 1 IPv4 VIP: 192.168.228.33 8SCAN VIP is enabled. 9SCAN VIP is individually enabled on nodes: 10SCAN VIP is individually disabled on nodes: 11SCAN 2 IPv4 VIP: 192.168.228.35 12SCAN VIP is enabled. 13SCAN VIP is individually enabled on nodes: 14SCAN VIP is individually disabled on nodes: 15SCAN 3 IPv4 VIP: 192.168.228.34 16SCAN VIP is enabled. 17SCAN VIP is individually enabled on nodes: 18SCAN VIP is individually disabled on nodes: Oracle Instance resources: 1# As oracle user 2srvctl config database 3srvctl config database -d <SID> 4srvctl status database -d <SID> 5srvctl status nodeapps -n host-env-datad1n1 6srvctl config nodeapps -n host-env-datad1n1 7# ============ 8srvctl stop database -d DB_NAME 9srvctl stop database -d DB_NAME -o normal 10srvctl stop database -d DB_NAME -o immediate 11srvctl stop database -d DB_NAME -o transactional 12srvctl stop database -d DB_NAME -o abort 13srvctl stop instance -d DB_NAME -i INSTANCE_NAME 14# ============= 15srvctl start database -d DB_NAME -n host-env-datad1n1 16srvctl start database -d DB_NAME -o nomount 17srvctl start database -d DB_NAME -o mount 18srvctl start database -d DB_NAME -o open 19# ============ 20srvctl relocate database -db DB_NAME -node host-env-datad1n1 21srvctl modify database -d DB_NAME -instance DB_NAME 22srvctl restart database -d DB_NAME 23# === Do not do it 24srvctl modify instance -db DB_NAME -instance DB_NAME_2 -node host-env-datad1n2 25srvctl modify database -d DB_NAME -instance DB_NAME 26srvctl modify database -d oraclath -instance oraclath Cluster resources 1crs_stat 2crsctl status res 3crsctl status res -t 4crsctl check cluster -all 5 6# Example how it should look: 7/opt/oracle/grid/12.2.0.1/bin/crsctl check cluster -all 8************************************************************** 9host-env-datad1n1: 10CRS-4535: Cannot communicate with Cluster Ready Services 11CRS-4529: Cluster Synchronization Services is online 12CRS-4534: Cannot communicate with Event Manager 13************************************************************** 14host-env-datad1n2: 15CRS-4537: Cluster Ready Services is online 16CRS-4529: Cluster Synchronization Services is online 17CRS-4533: Event Manager is online 18************************************************************** 1show parameter cluster 2 3NAME TYPE VALUE 4------------------------------------ ----------- ------------------------------ 5cdb_cluster boolean FALSE 6cdb_cluster_name string DB_NAME 7cluster_database boolean TRUE 8cluster_database_instances integer 2 9cluster_interconnects string Stop/start secondary node: 1-- Prevent Database to switch over 2ALTER database cluster_database=FALSE; 1# as root 2/u01/oracle/base/product/19.0.0/grid/bin/crsctl stop crs -f 3/u01/oracle/base/product/19.0.0/grid/bin/crsctl disable crs 4 5# Shutdown/startup VM or other actions 6 7# as root 8/u01/oracle/base/product/19.0.0/grid/bin/crsctl enable crs 9/u01/oracle/base/product/19.0.0/grid/bin/crsctl start crs Stop/Start properly DB on both nodes: 1# as oracle user 2srvctl stop database -d oraclath 3 4# As root user, on both nodes: 5/opt/oracle/grid/12.2.0.1/bin/crsctl stop crs -f 6/opt/oracle/grid/12.2.0.1/bin/crsctl disable crs 7 8# As root user, on both nodes: 9/opt/oracle/grid/12.2.0.1/bin/crsctl enable crs 10/opt/oracle/grid/12.2.0.1/bin/crsctl start crs 11 12# checks after restart 13ps -ef | grep asm_pmon | grep -v "grep" 14 15# if ASM is up and running 16srvctl start database -d oraclath -node host1-env-data1n1.domain Listner issue 1# As oracle user 2srvctl status scan_listener 3 4PRCR-1068 : Failed to query resources 5CRS-0184 : Cannot communicate with the CRS daemon. the solution:
🌌 How I Created This Blog
πŸ—Ώ Partition
πŸ—Ώ Partition
Checks your disks 1# check partion 2parted -l /dev/sda 3fdisk -l 4 5# check partition - visible before the mkfs 6ls /sys/sda/sda* 7ls /dev/sd* 8 9# give partition after the mkfs or pvcreate 10blkid 11blkid -o list 12 13# summary about the disks, partitions, FS and LVM 14lsblk 15lsblk -f Create Partition 1 on disk sdb in script mode 1# with fdisk 2printf "n\np\n1\n\n\nt\n8e\nw\n" | sudo fdisk "/dev/sdb" 3 4# with parted 5sudo parted /dev/sdb mklabel gpt mkpart primary 1 100% set 1 lvm on Gparted : interface graphique (ce base sur parted un utilitaire GNU - Table GPT)
🌱 MDadm
🌱 MDadm
The Basics mdadm (multiple devices admin) is software solution to manage RAID. It allow: create, manage, monitor your disks in an RAID array. you can the full disks (/dev/sdb, /dev/sdc) or (/dev/sdb1, /dev/sdc1) replace or complete raidtools Checks Basic checks 1# View real-time information about your md devices 2cat /proc/mdstat 3 4# Monitor for failed disks (indicated by "(F)" next to the disk) 5watch cat /proc/mdstat Checks RAID 1# Display details about the RAID array (replace /dev/md0 with your array) 2mdadm --detail /dev/md0 3 4# Examine RAID disks for information (not volume) similar to --detail 5mdadm --examine /dev/sd* Settings The conf file /etc/mdadm.conf does not exist by default and need to be created once you finish your install. This file is required for the autobuild at boot.
πŸ“‚ Filesystem
πŸ“‚ Filesystem
FS Types ext4 : the most widespread on GNU/Linux (derived from ext2 and ext3). It is journaled, meaning it records write operations to guarantee data integrity in case of an abrupt disk stop. It can also handle volumes up to 1 EiB (1024 PiB), and allows pre-allocating a contiguous area for a file to minimize fragmentation. Use this filesystem if you want to be able to read data back from macOS or Windows.
πŸ§ͺ SMART
πŸ§ͺ SMART
S.M.A.R.T. is a technology that allows you to monitor and analyze the health and performance of your hard drives. It provides valuable information about the status of your storage devices. Here are some useful commands and tips for using S.M.A.R.T. with smartctl: Display S.M.A.R.T. Information To display S.M.A.R.T. information for a specific drive, you can use the following command: 1smartctl -a /dev/sda This command will show all available S.M.A.R.T. data for the /dev/sda drive.
🧱 ISCSI
🧱 ISCSI
Install 1yum install iscsi-initiator-utils 2 3#Checks 4iscsiadm -m session -P 0 # get the target name 5iscsiadm -m session -P 3 | grep "Target: iqn\|Attached scsi disk\|Current Portal" 6 7# Discover and mount ISCSI disk 8iscsiadm -m discovery -t st -p 192.168.1.112 9iscsiadm --mode discovery --type sendtargets --portal 192.168.1.112 10 11# Login 12iscsiadm -m node -T iqn.1992-04.com.emc:cx.ckm00192201413.b0 -l 13iscsiadm -m node -T iqn.1992-04.com.emc:cx.ckm00192201413.b1 -l 14iscsiadm -m node -T iqn.1992-04.com.emc:cx.ckm00192201413.a1 -l 15iscsiadm -m node -T iqn.1992-04.com.emc:cx.ckm00192201413.a0 -l 16 17# Enable/Start service 18systemctl enable iscsid iscsi && systemctl stop iscsid iscsi && systemctl start iscsid iscsi Rescan BUS 1for BUS in /sys/class/scsi_host/host*/scan; do echo "- - -" > ${BUS} ; done 2 3sudo sh -c 'for BUS in /sys/class/scsi_host/host*/scan; do echo "- - -" > ${BUS} ; done ' Partition your FS
🩺 multipath
🩺 multipath
Install and Set Multipath 1yum install device-mapper-multipath Check settings in vim /etc/multipath.conf: 1defaults { 2user_friendly_names yes 3path_grouping_policy multibus 4} add disk in blacklisted and a block 1multipaths { 2 multipath { 3 wwid "36000d310004142000000000000000f23" 4 alias oralog1 5 } Special config for some providers. For example, recommended settings for all Clariion/VNX/Unity class arrays that support ALUA: 1 devices { 2 device { 3 vendor "DGC" 4 product ".*" 5 product_blacklist "LUNZ" 6 : 7 path_checker emc_clariion ### Rev 47 alua 8 hardware_handler "1 alua" ### modified for alua 9 prio alua ### modified for alua 10 : 11 } 12 } Checks config with: multipathd show config |more
🧐 LVM
🧐 LVM
The Basics list of component: PV (Physical Volume) VG (Volume Group) LV (Logical Volume) PE (Physical Extend) LE (Logical Extend) FS (File Sytem) LVM2 use a new driver, the device-mapper allow the us of diskΒ΄s sectors in different targets: - linear (most used in LVM). - stripped (stripped on several disks) - error (all I/O are consider in errors) - snapshot (allow snapshot async) mirror (integrate elements useful for the pvmove command) below example show you a striped volume and linear volume 1lvs --all --segments -o +devices 2server_xplore_col1 vgdata -wi-ao---- 21 striped 1.07t /dev/md2(40229),/dev/md3(40229),/dev/md4(40229),/dev/md5(40229),… 3server_xplore_col2 vgdata -wi-ao---- 1 linear 219.87g /dev/md48(0) Basic checks 1# Summary 2pvs 3vgs 4lvs 5 6# Scanner 7pvscan 8vgscan 9lvscan 10 11# Details info 12pvdisplay [sda] 13pvdisplay -m /dev/emcpowerd1 14vgdisplay [vg_root] 15lvdisplay [/dev/vg_root/lv_usr] 16 17# Summary details 18lvmdiskscan 19 /dev/sda1 [ 600.00 MiB] 20 /dev/sda2 [ 1.00 GiB] 21 /dev/sda3 [ 38.30 GiB] LVM physical volume 22 /dev/sdb1 [ <100.00 GiB] LVM physical volume 23 /dev/sdc1 [ <50.00 GiB] LVM physical volume 24 /dev/sdj [ 20.00 GiB] 25 1 disk 26 2 partitions 27 0 LVM physical volume whole disks 28 3 LVM physical volumes Usual Scenario in LVM Extend an existing LVM filesystem: 1parted /dev/sda resizepart 3 100% 2udevadm settle 3pvresize /dev/sda3 4 5# Extend a XFS to a fixe size 6lvextend -L 30G /dev/vg00/var 7xfs_growfs /dev/vg00/var 8 9# Add some space to a ext4 FS 10lvextend -L +10G /dev/vg00/var 11resize2fs /dev/vg00/var 12 13# Extend to a pourcentage and resize automaticly whatever is the FS type. 14lvextend -l +100%FREE /dev/vg00/var -r Create a new LVM filesystem: 1parted /dev/sdb mklabel gpt mkpart primary 1 100% set 1 lvm on 2udevadm settle 3pvcreate /dev/sdb1 4vgcreate vg01 /dev/sdb1 5lvcreate -n lv_data -l 100%FREE vg01 6 7# Create a XFS 8mkfs.xfs /dev/vg01/lv_data 9mkdir /data 10echo "/dev/mapper/vg01-lv_data /data xfs defaults 0 0" >> /etc/fstab 11mount -a 12 13# Create an ext4 14mkfs.ext4 /dev/vg01/lv_data 15mkdir /data 16echo "/dev/mapper/vg01-lv_data /data ext4 defaults 0 0" >> /etc/fstab 17mount -a Remove SWAP: 1swapoff -v /dev/dm-1 2lvremove /dev/vg00/swap 3vi /etc/fstab 4vi /etc/default/grub 5grub2-mkconfig -o /boot/efi/EFI/redhat/grub.cfg 6grubby --remove-args "rd.lvm.lv=vg00/swap" --update-kernel /boot/vmlinuz-3.10.0-1160.71.1.el7.x86_64 7grubby --remove-args "rd.lvm.lv=vg00swap" --update-kernel /boot/vmlinuz-3.10.0-1160.el7.x86_64 8grubby --remove-args "rd.lvm.lv=vg00/swap" --update-kernel /boot/vmlinuz-0-rescue-cd2525c8417d4f798a7e6c371121ef34 9echo "vm.swappiness = 0" >> /etc/sysctl.conf 10sysctl -p Move data form disk to another: 1# #n case of crash, just relaunch pvmove without arguments 2pvmove /dev/emcpowerd1 /dev/emcpowerc1 3 4# Remove PV from a VG 5vgreduce /dev/emcpowerd1 vg01 6 7# Remove all unused PV from VG01 8vgreduce -a vg01 9 10# remove all PV 11pvremove /dev/emcpowerd1 mount /var even if doesn’t want: 1lvchange -ay --ignorelockingfailure --sysinit vgroot/var Renaming: 1# VG rename 2vgrename 3 4# LV rename 5lvrename 6 7# PV does not need to be rename LVM on partition VS on Raw Disk Even if in the past I was using partition MS-DOS disklabel or GPT disklabel for PV, I prefer now to use directly LVM on the main block device. There is no reason to use 2 disklabels, unless you have a very specific use case (like disk with boot sector and boot partition).
πŸ› NFS
πŸ› NFS
The Basics NFS vs iscsi NFS can handle simultaniously writing from several clients. NFS is a filesystem , iscsi is a block storage. iscsi performance are same with NFS. iscsi will appear as disk to the OS, not the case for NFS. Concurrent access to a block device like iSCSI is not possible with standard file systems. You’ll need a shared disk filesystem (like GFS or OCSFS) to allow this, but in most cases the easiest solution would be to just use a network share (via SMB/CIFS or NFS) if this is sufficient for your application.
πŸ”οΈ Investigate
πŸ”οΈ Investigate
Ressources 1# in crontab or tmux session - take every hour a track of the memory usage 2for i in {1..24} ; do echo -n "===================== " ; date ; free -m ; top -b -n1 | head -n 15 ; sleep 3600; done >> /var/log/SYSADM/memory.log &
🚩 Compare
🚩 Compare
Compare files 1diff <file1> <file2> # -w to ignore whitespace. 2colordiff <file1> <file2> # colourised diff. 3wdiff <file1> <file2> # word diff: [βˆ’ βˆ’] replaced word, {+ +} added word. 4vimdiff <file1> <file2> # open both files in vim (blue = entirely different lines, red = partially different). 5fgrep -f <list> <file> # compare two lists (e.g. of hosts). Compare jar files 1diff -W200 -y <(unzip -vqq file1.jar | awk '{ if ($1 > 0) {printf("%s\t%s\n", $1, $8)}}' | sort -k2) <(unzip -vqq file2.jar | awk '{ if ($1 > 0) {printf("%s\t%s\n", $1, $8)}}' | sort -k2)
🚩 Files
🚩 Files
Find a process blocking a file with fuser: 1fuser -m </dir or /files> # Find process blocking/using this directory or files. 2fuser -cu </dir or /files> # Same as above but add the user 3fuser -kcu </dir or /files> # Kill process 4fuser -v -k -HUP -i ./ # Send HUP signal to process 5 6# Output will send you <PID + letter>, here is the meaning: 7# c current directory. 8# e executable being run. 9# f open file. (omitted in default display mode). 10# F open file for writing. (omitted in default display mode). 11# r root directory. 12# m mmap'ed file or shared library. with lsof ( = list open file): 1lsof +D /var/log # Find all files blocked with the process and user. 2lsof -a +L1 <mountpoint> # Process blocking a FS. 3lsof -c ssh -c init # Find files open by thoses processes. 4lsof -p 1753 # Find files open by PID process. 5lsof -u root # Find files open by user. 6lsof -u ^user # Find files open by user except this one. 7kill -9 `lsof -t -u toto` # kill user's processes. (option -t output only PID). MacGyver method: 1#When you have no fuser or lsof: 2find /proc/*/fd -type f -links 0 -exec ls -lrt {} \; AIX specifics (fuser) 1fuser -d /tmp # see the processes using the /tmp directory (AIX) 2fuser -c /your_FS # all processes with an open file in the filesystem (AIX) 3fuser -cu /dev/vg01/lvol5 # also search with a filesystem or an LV -c == -m ; -u also shows the process user. to kill the processes: fuser -kcu. File deleted but space still held For detecting deleted-but-still-open files (lsof +L1) and freeing the held space, see the Disk Cleanup page.
🚩 Network Manager
🚩 Network Manager
Basic Troubleshooting Checks interfaces 1nmcli con show 2NAME UUID TYPE DEVICE 3ens192 4d0087a0-740a-4356-8d9e-f58b63fd180c ethernet ens192 4ens224 3dcb022b-62a2-4632-8b69-ab68e1901e3b ethernet ens224 5 6nmcli dev status 7DEVICE TYPE STATE CONNECTION 8ens192 ethernet connected ens192 9ens224 ethernet connected ens224 10ens256 ethernet connected ens256 11lo loopback unmanaged -- 12 13# Get interfaces details : 14nmcli connection show ens192 15nmcli -p con show ens192 16 17# Get DNS settings in interface 18UUID=$(nmcli --get-values connection.uuid c show "cloud-init eth0") 19nmcli --get-values ipv4.dns c show $UUID Changing Interface name 1nmcli connection add type ethernet mac "00:50:56:80:11:ff" ifname "ens224" 2nmcli connection add type ethernet mac "00:50:56:80:8a:0b" ifname "ens256" Create a custom config 1nmcli con load /etc/sysconfig/network-scripts/ifcfg-ens224 2nmcli con up ens192 Adding a Virtual IP 1nmcli con mod enp1s0 +ipv4.addresses "192.168.122.11/24" 2ip addr del 10.10.10.36/24 dev ens160 3 4nmcli con reload # before to reapply 5nmcli device reapply ens224 6systemctl status network.service 7systemctl restart network.service Add a DNS entry 1UUID=$(nmcli --get-values connection.uuid c show "cloud-init eth0") 2DNS_LIST=$(nmcli --get-values ipv4.dns c show $UUID) 3nmcli conn modify "$UUID" ipv4.dns "${DNS_LIST} ${DNS_IP}" 4 5# /etc/resolved is managed by systemd-resolved 6sudo systemctl restart systemd-resolved
🎢 Samba / CIFS
🎢 Samba / CIFS
Server Side First Install samba and samba-client (for debug + test) /etc/samba/smb.conf 1[home] 2Workgroup=WORKGROUP (le grp par defaul sur windows) 3Hosts allow = ... 4[shared] 5browseable = yes 6path = /shared 7valid users = user01, @un_group_au_choix 8writable = yes 9passdb backend = tdbsam #passwords are stored in the /var/lib/samba/private/passdb.tdb file. Test samba config testparm /usr/bin/testparm -s /etc/samba/smb.conf smbclient -L \192.168.56.102 -U test : list all samba shares available smbclient //192.168.56.102/sharedrepo -U test : connect to the share pdbedit -L : list user smb (better than smbclient)
🍻 SSHFS
🍻 SSHFS
SSHFS SSHFS mounts a remote filesystem on your local filesystem through an SSH connection, all with user rights. The advantage is being able to manipulate remote data with any file manager (Nautilus, Konqueror, ROX, or even the command line). - Prerequisites: administrator rights, ethernet connection, installation of FUSE and the SSHFS package. - SSHFS users must belong to the `fuse` group. Note: FUSE allows a user to mount a filesystem themselves. Normally, mounting a filesystem requires being an administrator, or having it pre-approved in /etc/fstab with hard-coded information.
πŸ‘€ Users & Connections
πŸ‘€ Users & Connections
Investigate a user 1last # the last user connections to a server (based on /var/log/wtmp or btmp). 2ac -d # statistics of my connection time per day. 3ac -p <user> # the connection time of all users (or of a specific user). 4finger # who is connected (-l to also see mails and plans of all users). 5w # who is connected, doing what, and how much CPU they use. 6who # who is connected (-u for more info: PID, etc.). 7who am i # with which login I am connected. 8id -a # all info about the user I'm connected as (more precise than "who am i"). 9logname # the login name of the current account. Reboots & uptime 1last reboot # see all the reboots that took place. 2uptime # see how long the server has been up + the load average. 3lslogins -L # also shows whether a user shutdown/rebooted the machine.
πŸ“œ Logs
πŸ“œ Logs
Where the system logs live On a Unix machine, the system logs are in /var/log/messages (or /var/adm/messages on Solaris). This is where you find the errors, with log rotation. Default syslog output Linux Solaris HP-UX AIX BSD location /var/log/messages, /var/log/secure, /var/log/boot.log /var/adm/messages /var/adm/syslog/mail.log, /var/adm/syslog/syslog.log /tmp or none /var/log/syslog System accounting (login & process) Type Linux Solaris HP-UX AIX current logins /var/run/utmp /var/adm/utmpx /var/adm/utmp /etc/utmp login history /var/log/wtmp /var/adm/wtmpx /var/adm/wtmp /var/adm/wtmp process accounting /var/log/pacct /var/adm/pacct /var/adm/pacct /var/adm/pacct Login errors Linux Solaris HP-UX AIX failed logins /var/log/btmp, /var/log/messages /var/adm/loginlog, /var/adm/sulog /var/adm/sulog /etc/security/failedlogin Investigate the logs 1# today's logs 2grep "$(date '+%b %d')" /var/log/messages 3 4# disk errors (nawk: print the last field of the "Error Block" lines) 5nawk '/Error Block/{print $NF}' /var/adm/messages* | sort | uniq 6 7# find the IPs in a log, sort them and remove the duplicates 8cat /var/log/maillog | grep -Eo '([0-9]{1,3}\.){3}[0-9]{1,3}' | sort -n -t . -k 1,1 -k 2,2 -k 3,3 -k 4,4 | uniq Network investigation 1# ping a list of servers 2for ip in $(awk '/192.168.45/ {print $1}' /etc/hosts); do ping -c 1 $ip; done Investigate on several servers 1for vm in vm{1..27}; do ssh -q $vm "hostname; free; sar -r 3 3"; done SSH authentication logs /var/log/auth.log β€” SSH connection logs. Check that there are not too many failed connections (a sign of an intrusion attempt).
πŸ”Ž Search, Find & Compare
πŸ”Ž Search, Find & Compare
Find files quickly 1locate <pattern> # find a directory or file quickly (uses an index); a brand-new file won't be found. 2updatedb # update the locate index. Open a file 1view <file> # opens a read-only vi view (preferred if you just want to search/view). 2gzcat / zcat <file.gz> # read a gzipped file. Info on a file or directory 1stat </my/file> # all info about a file (inode, creation, modification, access dates, etc.). 2stat -f <FS> # info about a filesystem. 3stat -c%s $LOGFILE # [scripting] get a precise value (size, modification date, etc.). grep 1grep -w 'xyz' # match the whole word. 2grep -x 'Hello, world!' # the whole line must match. 3grep -c <pattern> # count the matching lines. 4grep -l "ERROR:" *.log # search all .log files, list the files that match. 5grep -L <pattern> # inverse: list the files that do NOT match. 6grep -f <patternfile> <file> # apply the patterns read from patternfile. 7grep -i <pattern> # ignore case. 8grep -v <pattern> # return the lines that do NOT match. 9grep -m x <pattern> # stop after x matching lines. 10grep -n <pattern> # show the line number. 11grep -q <pattern> # quiet: exit 0 if found, 1 (or 2) otherwise (for scripting). 12grep -s <pattern> # suppress permission/inexistent-file error messages. 13grep -H <pattern> # show the filename next to each matching line. 14grep -h <pattern> # do not show the filename (default behaviour). 15grep -A x <pattern> # also show x lines After. 16grep -B x <pattern> # also show x lines Before. 17grep -C x <pattern> # show x lines of context (A + B). 18grep -a <pattern> <binary> # search a binary file as if it were text. 1egrep = grep -E # for complex regular expressions. 1# extract the 3rd field, then cut: 2cat file | grep /u01/grid/19c | awk '{print $3}' | cut -f2 -d'"' 3# is equivalent to: 4cat file | grep -o /u01/grid/19c
⏲️ Cron & Anacron
⏲️ Cron & Anacron
Configurations /etc/crontab : the daemon’s configuration file, defining the default behaviour of crond (SHELL, MAILTO, etc.). /etc/cron.d/... : system crontab (used by admins). /var/spool/cron/root : crontab per user. “Day of month” and “Day of week” are combined with a logical OR, so if both are set, the job runs on that day of the month and on that day of the week. 1MAILTO="admin@example.com" 2* * * * * root /usr/local/sbin/mycommand.sh > /dev/null 2>&1 Anacron /etc/anacrontab : file that runs, via the run-parts command, /etc/cron.daily, /etc/cron.weekly, /etc/cron.monthly. /etc/cron.d/0hourly : exception, runs /etc/cron.hourly via run-parts, checking the last run of the task in /var/spool/anacron/.... Special cases Exactly the last day of each month:
πŸ“œ Logrotate
πŸ“œ Logrotate
1logrotate -d /etc/logrotate.d/app # test a new configuration. 2reading config info for /data/log/app 3Handling 1 logs 4rotating pattern: /data/log/app after 1 days (10 rotations) 5empty log files are rotated, old logs are removed Options Usage: logrotate [OPTION...] <configfile> -d, --debug Don't do anything, just test (implies -v) -f, --force Force file rotation -m, --mail=command Command to send mail (instead of `/bin/mail') -s, --state=statefile Path of state file -v, --verbose Display messages during rotation
πŸ“¦ Chroot Jail
πŸ“¦ Chroot Jail
Change the root directory of a command or a process, and its children. In no case should `chroot` be relied upon as a security boundary β€” a process running as root can escape the jail. Example: creating a chroot 1# create the "jail" directory 2J=$HOME/jail 3mkdir -p $J 4mkdir -p $J/{bin,lib64,lib} 5cd $J 6 7# copy the binaries and their libraries into the jail 8cp -v /bin/{bash,ls} $J/bin 9 10list="$(ldd /bin/bash | egrep -o '/lib.*\.[0-9]')" 11for i in $list; do cp -v "$i" "${J}${i}"; done 12 13list="$(ldd /bin/ls | egrep -o '/lib.*\.[0-9]')" 14for i in $list; do cp -v "$i" "${J}${i}"; done 15 16# enter the jail 17sudo chroot $J /bin/bash
πŸ” Runlevels & Shutdown
πŸ” Runlevels & Shutdown
Shutdown / reboot Solaris Red Hat Ubuntu / Debian HP-UX AIX Power down shutdown -i5 -g0 -y shutdown -h shutdown -h shutdown -h now shutdown -F Reboot shutdown -i6 -g0 -y shutdown -r shutdown -r shutdown -r now shutdown -Fr OK prompt shutdown -i0 -g0 -y β€” β€” β€” β€” Fast reboot -- -r (reconfigure) shutdown -f (no fsck) shutdown -P (power off) shutdown -F (force fsck) β€” Force fsck touch /reconfigure touch /forcefsck edit /etc/default/rcS β†’ FSCKFIX=yes β€” β€” Change runlevel Tool Solaris Red Hat Ubuntu / Debian HP-UX AIX halt βœ… βœ… βœ… βœ… βœ… init βœ… βœ… βœ… βœ… βœ… poweroff βœ… βœ… βœ… βœ… βœ… reboot βœ… βœ… βœ… βœ… βœ… shutdown βœ… βœ… βœ… βœ… βœ… telinit βœ… βœ… βœ… β€” βœ… uadmin βœ… β€” β€” β€” β€” Runlevels Level Solaris Red Hat Ubuntu / Debian HP-UX AIX 0 shutdown halt halt halt reserved 1 single user single user single user single user reserved 2 n/a multiuser (no networking) multiuser (default) multiuser (networking) multiuser + NFS 3 multi-user multiuser (networking) same as 2 multiuser + NFS + CDE GUI (default) user defined 4 n/a unused same as 2 multiuser + NFS + VUE GUI user defined 5 power off GUI same as 2 n/a user defined 6 reboot reboot reboot n/a user defined 7-9 β€” β€” β€” β€” user defined Change the default runlevel Solaris / Red Hat / HP-UX / AIX : edit the initdefault line in vi /etc/inittab. Ubuntu / Debian : edit vi /etc/event.d/rc-default. On systemd systems (RHEL 7+, Ubuntu 15+), the SysV runlevels are replaced by targets β€” e.g. systemctl isolate multi-user.target (runlevel 3), systemctl isolate graphical.target (runlevel 5), systemctl set-default multi-user.target. See the Systemd page.
πŸ• NTP & Time Synchronisation
πŸ• NTP & Time Synchronisation
Client verification (ntpd / chrony) 1ntpstat # see which NTP server we synchronise with, and whether the sync is good. 1synchronised to NTP server (192.168.1.12) at stratum 4 2 time correct to within 68 ms 3 polling server every 1024 s 1ntpq -p # see the state of the peers. 2ntpq -c peers remote refid st t when poll reach delay offset jitter ============================================================================== +192.168.1.11 192.168.2.4 4 u 259 1024 373 0.731 -0.980 0.557 *192.168.1.12 192.168.3.21 3 u 385 1024 377 0.773 0.146 0.365 192.168.4.255 .BCST. 16 u - 64 0 0.000 0.000 0.000 The server preceded by an asterisk (*) is the one being used. Those preceded by a - are currently discarded by the server-selection algorithm. Those whose name is preceded by a + are possible synchronisation candidates. A server preceded by a space is either unreachable or too distant. Column meaning remote β€” the server name. refid β€” the parent server’s identifier. st β€” the server’s stratum. t β€” the server type. when β€” seconds elapsed since the last contact. poll β€” seconds between each contact. reach β€” bitmask of successful contacts (octal): the server considers itself synchronised when reach reaches 177; a quality, stable connection shows 377. delay β€” estimated round-trip time (ms) of the UDP packet. offset β€” estimated difference between the peer’s clock and the internal clock. jitter β€” dispersion of the reference values obtained from this peer. Restart the NTP daemon 1service ntpd restart # or: systemctl restart ntpd Configuration & logs 1cat /etc/ntp.conf # "server example.com" + restart ntpd + enable 2/var/log/ntpstats ntpdate (legacy) Old service that synchronises NTP at boot (install the package first).
πŸ‘₯ Users & Groups
πŸ‘₯ Users & Groups
Configuration files File Check command Purpose /etc/passwd pwck user accounts /etc/group grpck groups /etc/shadow β€” password hashes and aging /etc/gshadow β€” group passwords /etc/skel β€” files installed by default when a user is created Basic commands 1useradd -g <GID> -G <GID2> <user> # create a user in primary group GID (and supplementary group GID2). 2usermod <options> <user> # modify a user. 3userdel -r <user> # delete a user (and its home directory). 4groupadd / groupmod / groupdel # manage groups. 5 6id -a # show all info about the current user (UID, GUID, groups, etc.) - more precise than "who am i". 7sg <group> -c '<command>' # execute a command as a different group ID (to run scripts or write to a file with group rights). Password management 1passwd -u <user> # unlock a user account. 2echo "password" | passwd --stdin <user> # scripted password change. Account aging 1chage -l <user> # see the expiration dates. 1# list the expiry of every account 2for account in $(cut -f1 -d: /etc/passwd); do 3 echo "ACCOUNT: $account , EXPIRES: $(chage -l $account | grep 'Account expires' | awk '{print $4, $5, $6}'), CHANGED: $(chage -l $account | grep 'Last password change' | awk '{print $5, $6, $7}')"; 4done 1# change the aging info interactively 2chage <user> To unlock an account, set “Last Password Change” to -1 in chage (or use passwd -u <user>).