Video summary

2 ROOTS 1 CAP | Linux Capabilities под капотом

Main summary

Key takeaways

Educational

Main ideas / lessons

  • “Being root” in Linux is not automatically “all-powerful.” Root processes can still be restricted depending on which Linux capabilities are enabled in their credentials/effective set.

  • Linux shifted from the old Unix privilege model to a capability-based model (starting with Linux kernel 2.2). Instead of one binary “UID 0 = bypass everything,” capabilities let the kernel grant specific privileged powers (e.g., bind privileged ports, bypass certain access checks).

  • Capabilities are permission bits evaluated by the kernel. Even when two containers run as UID=0 and GID=0, they may differ in their effective capability bitmasks, leading to different outcomes (e.g., file read allowed vs denied).

  • The video demonstrates a concrete difference: only one capability bit changes the result. In the two containers, the only difference in the effective capability sets is:

    • Capability: CAP_DAC_OVERRIDE (shown via CDAC override in the transcript)

Result:

- Container with the bit **set** can read a file whose DAC permissions would normally deny it.
- Container with the bit **cleared** receives **“Permission denied.”**
  • Docker/Kubernetes don’t implement a special permission system for this. They configure the container’s Linux process credentials/capabilities; the kernel is still the final authority.

  • Security implication: minimize capabilities.

    • If an attacker gains code execution inside a container, they inherit the process’s capabilities.
    • A broader capability set (especially “admin-like” capabilities such as CAP_SYS_ADMIN) can dramatically increase attacker options.
    • Therefore, the recommended approach is typically:
      • Drop all capabilities
      • Add back only what’s necessary
    • Also use Kubernetes/Docker security hardening options (non-root, disallow privilege escalation, restricted syscalls, etc.).

Methodology / step-by-step approach shown in the video (Linux/Docker capability inspection)

  • Set up comparison containers

    • Run two containers from the same image on the same host.
    • In both:
      • the target file has the same owner/rights (owner root; rights = none for reading/writing/execute).
      • both processes run as UID=0 / GID=0.
    • In only one container:
      • remove one capability bit (specifically the one corresponding to DAC override).
  • Reproduce the difference

    • Container #1: attempt to read the “secret” file → succeeds.
    • Container #2: attempt to read the same file → Permission denied.
  • Inspect process capability state using /proc (no special Docker tools required)

    • Use /proc/<pid>/status and search for lines starting with Cap.
    • Focus on CapEff (effective capabilities).
    • Observe that:
      • most of the effective capability bitmask appears identical
      • but there is a difference in one hex digit/bit pattern at the end.
  • Compute what changed

    • Use XOR logic (bitwise exclusive-or) between the two hex masks to find which capability bits differ.
    • Convert the resulting XOR mask into binary to see which bit(s) differ.
    • Identify the differing capability by name using a “decode” tool:
      • the differing capability is CAP_DAC_OVERRIDE.
  • Verify capability ID mapping

    • Cross-check the capability’s numeric index (bit position) against Linux kernel headers (via grep for the definition, e.g. capability number/bit).
  • Connect capability to the observed behavior

    • Explain the order of checks for file open/access:
      • Kernel checks standard DAC permissions first (which deny reading due to no read bits set).
      • Then the kernel checks whether CAP_DAC_OVERRIDE is effective:
        • If present → bypass DAC denial → read succeeds.
        • If absent → DAC denial stands → Permission denied.

Docker / Kubernetes configuration concepts mentioned

Docker capability controls

  • --cap-add <cap>: add a capability
  • --cap-drop <cap>: remove a capability
  • --cap-drop ALL:
    • the process may still start with UID 0
    • but its original capability set is largely removed (and only those explicitly added would remain)

Caveat noted in transcript: “Dropping capabilities” is not the whole story of container privilege; it can also affect access to devices and other container restrictions.

Kubernetes / container runtime security controls (as described)

  • Kubernetes passes configuration to the container runtime, which then creates a normal Linux process and relies on the kernel for enforcement.

  • Example hardening strategy described (conceptually):

    • Drop all capabilities, then add only needed ones
    • Use security flags such as:
      • allowPrivilegeEscalation: false (control exec/privilege behavior)
      • runAsNonRoot: true (avoid UID 0 inside container)
      • Often also set runAsUser
    • Use an seccomp profile / system call filtering:
      • seccomp filters syscalls, reducing what a process can do even if it has some capabilities
    • Mention of runtime default security profile from the runtime itself (e.g., container runtime-provided profile).

Practical benefit / why this matters (security motivation)

  • If a containerized app is compromised via RCE, the attacker executes code in the context of an existing process.
  • The attacker inherits:
    • namespaces,
    • IDs,
    • and especially the process’s capabilities.
  • Therefore:
    • Two “root” containers can be radically different attacker starting points.
    • Capability grants expand the set of privileged operations available to the attacker.
  • Best practice:
    • Keep capabilities minimal to reduce post-exploitation options.

Speakers / sources featured

  • Speaker/Presenter: The video narrator/author (referred to as “Gaias” in the transcript; no external name clearly provided)

  • Software/Systems referenced as sources of concepts:

    • Linux kernel capability model (kernel documentation/header definitions; concept introduced circa Linux 2.2)
    • /proc pseudo-filesystem (for capability inspection: CapEff etc.)
    • Docker (capability configuration via cap add/drop/cap drop all)
    • Kubernetes (securityContext + runtime behavior)
    • VFS (Virtual File System) (kernel layer involved in file access checks)
    • Linux kernel headers (capability numeric/bit definitions)

Original video