Highlights
- Multi-service nginx routing, systemd units (Restart=always), NVMe provisioning, GPU/VRAM hygiene, S3 estate audits and cost analysis.
- Diagnosed ephemeral instance-store loss (about 90 GB gone) and instituted a code-plus-cache backup and transfer pipeline with SHA-256 manifests.
- Server-to-server migrations done as agent-to-agent handoffs: ports, environment, nginx and shared DB briefs for the destination.