Video summary
2 ROOTS 1 CAP | Linux Capabilities под капотом
Main summary
Key takeaways
Main ideas / lessons
-
“Being root” in Linux is not automatically “all-powerful.” Root processes can still be restricted depending on which Linux capabilities are enabled in their credentials/effective set.
-
Linux shifted from the old Unix privilege model to a capability-based model (starting with Linux kernel 2.2). Instead of one binary “UID 0 = bypass everything,” capabilities let the kernel grant specific privileged powers (e.g., bind privileged ports, bypass certain access checks).
-
Capabilities are permission bits evaluated by the kernel. Even when two containers run as UID=0 and GID=0, they may differ in their effective capability bitmasks, leading to different outcomes (e.g., file read allowed vs denied).
-
The video demonstrates a concrete difference: only one capability bit changes the result. In the two containers, the only difference in the effective capability sets is:
- Capability:
CAP_DAC_OVERRIDE(shown viaCDAC overridein the transcript)
- Capability:
Result:
- Container with the bit **set** can read a file whose DAC permissions would normally deny it.
- Container with the bit **cleared** receives **“Permission denied.”**
-
Docker/Kubernetes don’t implement a special permission system for this. They configure the container’s Linux process credentials/capabilities; the kernel is still the final authority.
-
Security implication: minimize capabilities.
- If an attacker gains code execution inside a container, they inherit the process’s capabilities.
- A broader capability set (especially “admin-like” capabilities such as
CAP_SYS_ADMIN) can dramatically increase attacker options. - Therefore, the recommended approach is typically:
- Drop all capabilities
- Add back only what’s necessary
- Also use Kubernetes/Docker security hardening options (non-root, disallow privilege escalation, restricted syscalls, etc.).
Methodology / step-by-step approach shown in the video (Linux/Docker capability inspection)
-
Set up comparison containers
- Run two containers from the same image on the same host.
- In both:
- the target file has the same owner/rights (owner root; rights = none for reading/writing/execute).
- both processes run as UID=0 / GID=0.
- In only one container:
- remove one capability bit (specifically the one corresponding to DAC override).
-
Reproduce the difference
- Container #1: attempt to read the “secret” file → succeeds.
- Container #2: attempt to read the same file → Permission denied.
-
Inspect process capability state using
/proc(no special Docker tools required)- Use
/proc/<pid>/statusand search for lines starting withCap. - Focus on
CapEff(effective capabilities). - Observe that:
- most of the effective capability bitmask appears identical
- but there is a difference in one hex digit/bit pattern at the end.
- Use
-
Compute what changed
- Use XOR logic (bitwise exclusive-or) between the two hex masks to find which capability bits differ.
- Convert the resulting XOR mask into binary to see which bit(s) differ.
- Identify the differing capability by name using a “decode” tool:
- the differing capability is
CAP_DAC_OVERRIDE.
- the differing capability is
-
Verify capability ID mapping
- Cross-check the capability’s numeric index (bit position) against Linux kernel headers (via
grepfor the definition, e.g. capability number/bit).
- Cross-check the capability’s numeric index (bit position) against Linux kernel headers (via
-
Connect capability to the observed behavior
- Explain the order of checks for file open/access:
- Kernel checks standard DAC permissions first (which deny reading due to no read bits set).
- Then the kernel checks whether
CAP_DAC_OVERRIDEis effective:- If present → bypass DAC denial → read succeeds.
- If absent → DAC denial stands → Permission denied.
- Explain the order of checks for file open/access:
Docker / Kubernetes configuration concepts mentioned
Docker capability controls
--cap-add <cap>: add a capability--cap-drop <cap>: remove a capability--cap-drop ALL:- the process may still start with UID 0
- but its original capability set is largely removed (and only those explicitly added would remain)
Caveat noted in transcript: “Dropping capabilities” is not the whole story of container privilege; it can also affect access to devices and other container restrictions.
Kubernetes / container runtime security controls (as described)
-
Kubernetes passes configuration to the container runtime, which then creates a normal Linux process and relies on the kernel for enforcement.
-
Example hardening strategy described (conceptually):
- Drop all capabilities, then add only needed ones
- Use security flags such as:
allowPrivilegeEscalation: false(control exec/privilege behavior)runAsNonRoot: true(avoid UID 0 inside container)- Often also set
runAsUser
- Use an seccomp profile / system call filtering:
seccompfilters syscalls, reducing what a process can do even if it has some capabilities
- Mention of runtime default security profile from the runtime itself (e.g., container runtime-provided profile).
Practical benefit / why this matters (security motivation)
- If a containerized app is compromised via RCE, the attacker executes code in the context of an existing process.
- The attacker inherits:
- namespaces,
- IDs,
- and especially the process’s capabilities.
- Therefore:
- Two “root” containers can be radically different attacker starting points.
- Capability grants expand the set of privileged operations available to the attacker.
- Best practice:
- Keep capabilities minimal to reduce post-exploitation options.
Speakers / sources featured
-
Speaker/Presenter: The video narrator/author (referred to as “Gaias” in the transcript; no external name clearly provided)
-
Software/Systems referenced as sources of concepts:
- Linux kernel capability model (kernel documentation/header definitions; concept introduced circa Linux 2.2)
/procpseudo-filesystem (for capability inspection:CapEffetc.)- Docker (capability configuration via cap add/drop/cap drop all)
- Kubernetes (securityContext + runtime behavior)
- VFS (Virtual File System) (kernel layer involved in file access checks)
- Linux kernel headers (capability numeric/bit definitions)