Most Neoclouds Suck At Security
OpenAI vs HuggingFace, Container Escapes, Kernel Bypass, Network Policies, Security Keys, Multi-tenant Grafana, and a ClusterMAX 3.0 Preview

TL;DR
- Neocloud CISOs are increasingly involved in security negotiations due to growing counterparty risks in AI infrastructure supply chains.
- The article details five frightening patterns of cybersec vulnerabilities observed in neocloud testing.
- Despite AI's capabilities, its direct impact on overall CVE statistics is not yet statistically significant, though specific areas like Nvidia and AMD AI stacks show growth.
- The OpenAI vs HuggingFace security incident serves as a case study of AI agents coordinating attacks.
- Common bad design patterns include containers/VMs as sole isolation, multi-tenant Kubernetes control planes, lack of network isolation, and improper InfiniBand key configuration.
- Container escapes and outdated software versions remain prevalent issues, with providers needing systems for continuous patching.
- BlueField DPUs can pose security risks if left in the default 'host-trusted' mode, allowing host access to the DPU's Arm cores.
- Closed-source AI models have guardrails that hinder whitehat researchers in building Proofs of Concept (POCs) for vulnerabilities.
- Neoclouds are urged to update software to minimum versions, fix bad designs (e.g., single points of failure, inadequate isolation), and implement automated security bulletin monitoring.
- The ClusterMAX CLI is offered as a tool for auditing security and identifying outdated software with known vulnerabilities.