summaryrefslogtreecommitdiffstats
path: root/Documentation/admin-guide/binfmt-misc.rst
diff options
context:
space:
mode:
authorKees Cook <kees+treewide@kernel.org>2026-09-02 15:31:14 -0700
committerKees Cook <kees@kernel.org>2026-09-04 21:37:00 -0700
commit3a2c4d55e32ad65efebdb6de44eef3bfa08bb49d (patch)
treec65086f9bdcd48c6360fb7cb4598bca084da1f32 /Documentation/admin-guide/binfmt-misc.rst
downloadlinux-stable-3a2c4d55e32ad65efebdb6de44eef3bfa08bb49d.tar.gz
linux-stable-3a2c4d55e32ad65efebdb6de44eef3bfa08bb49d.zip
treewide: refresh kmalloc_obj() conversionsgrafted
This is another run of the Coccinelle script for converting kmalloc() family of allocations to kmalloc_obj() via the existing rules in scripts/coccinelle/api/kmalloc_objs.cocci This catches both the set of kmalloc() uses added since the first kmalloc_obj() conversions in v7.0 and adds a large group missed in the first pass due to Coccinelle not interacting well with the cleanup.h scoped_...() family of macros[1]. I worked around this with spatch's "--macro-file" argument to a file with all the scoped_...() macros mapped to Coccinelle's YACFE_ITERATOR[2] as that was the closest viable control flow indicator I could find. Build tested allmodconfig on x86, arm64, arm, loongarch, mips, powerpc, riscv, and s390 with no new warnings. Link: https://lore.kernel.org/lkml/202609021314.8A9C0B8@keescook/ [1] Link: https://github.com/coccinelle/coccinelle/blob/master/standard.h [2] Signed-off-by: Kees Cook <kees+treewide@kernel.org>
Diffstat (limited to 'Documentation/admin-guide/binfmt-misc.rst')
-rw-r--r--Documentation/admin-guide/binfmt-misc.rst391
1 files changed, 391 insertions, 0 deletions
diff --git a/Documentation/admin-guide/binfmt-misc.rst b/Documentation/admin-guide/binfmt-misc.rst
new file mode 100644
index 000000000..d26b63a27
--- /dev/null
+++ b/Documentation/admin-guide/binfmt-misc.rst
@@ -0,0 +1,391 @@
+Kernel Support for miscellaneous Binary Formats (binfmt_misc)
+=============================================================
+
+This Kernel feature allows you to invoke almost (for restrictions see below)
+every program by simply typing its name in the shell.
+This includes for example compiled Java(TM), Python or Emacs programs.
+
+To achieve this you must tell binfmt_misc which interpreter has to be invoked
+with which binary. Binfmt_misc recognises the binary-type by matching some bytes
+at the beginning of the file with a magic byte sequence (masking out specified
+bits) you have supplied. Binfmt_misc can also recognise a filename extension
+aka ``.com`` or ``.exe``.
+
+First you must mount binfmt_misc::
+
+ mount binfmt_misc -t binfmt_misc /proc/sys/fs/binfmt_misc
+
+To actually register a new binary type, you have to set up a string looking like
+``:name:type:offset:magic:mask:interpreter:flags`` (where you can choose the
+``:`` upon your needs) and echo it to ``/proc/sys/fs/binfmt_misc/register``.
+
+Here is what the fields mean:
+
+- ``name``
+ is an identifier string. A new /proc file will be created with this
+ name below ``/proc/sys/fs/binfmt_misc``; cannot contain slashes ``/`` for
+ obvious reasons.
+- ``type``
+ is the type of recognition. Give ``M`` for magic, ``E`` for extension and
+ ``B`` for a bpf-backed handler (see below).
+- ``offset``
+ is the offset of the magic/mask in the file, counted in bytes. This
+ defaults to 0 if you omit it (i.e. you write ``:name:type::magic...``).
+ Ignored when using filename extension matching.
+- ``magic``
+ is the byte sequence binfmt_misc is matching for. The magic string
+ may contain hex-encoded characters like ``\x0a`` or ``\xA4``. Note that you
+ must escape any NUL bytes; parsing halts at the first one. In a shell
+ environment you might have to write ``\\x0a`` to prevent the shell from
+ eating your ``\``.
+ If you chose filename extension matching, this is the extension to be
+ recognised (without the ``.``, the ``\x0a`` specials are not allowed).
+ Extension matching is case sensitive, and slashes ``/`` are not allowed!
+- ``mask``
+ is an (optional, defaults to all 0xff) mask. You can mask out some
+ bits from matching by supplying a string like magic and as long as magic.
+ The mask is anded with the byte sequence of the file. Note that you must
+ escape any NUL bytes; parsing halts at the first one. Ignored when using
+ filename extension matching.
+- ``interpreter``
+ is the program that should be invoked with the binary as first
+ argument (specify the full path). For ``B`` entries this field
+ carries the name of the bpf handler instead (see below).
+- ``flags``
+ is an optional field that controls several aspects of the invocation
+ of the interpreter. It is a string of capital letters, each controls a
+ certain aspect. The following flags are supported:
+
+ ``P`` - preserve-argv[0]
+ Legacy behavior of binfmt_misc is to overwrite
+ the original argv[0] with the full path to the binary. When this
+ flag is included, binfmt_misc will add an argument to the argument
+ vector for this purpose, thus preserving the original ``argv[0]``.
+ e.g. If your interp is set to ``/bin/foo`` and you run ``blah``
+ (which is in ``/usr/local/bin``), then the kernel will execute
+ ``/bin/foo`` with ``argv[]`` set to ``["/bin/foo", "/usr/local/bin/blah", "blah"]``. The interp has to be aware of this so it can
+ execute ``/usr/local/bin/blah``
+ with ``argv[]`` set to ``["blah"]``.
+ ``O`` - open-binary
+ Legacy behavior of binfmt_misc is to pass the full path
+ of the binary to the interpreter as an argument. When this flag is
+ included, binfmt_misc will open the file for reading and pass its
+ descriptor into the auxilary vector with the key "AT_EXECFD", thus
+ allowing the interpreter to execute non-readable binaries. This
+ feature should be used with care - the interpreter has to be trusted
+ not to emit the contents of the non-readable binary.
+ ``C`` - credentials
+ Currently, the behavior of binfmt_misc is to calculate
+ the credentials and security token of the new process according to
+ the interpreter. When this flag is included, these attributes are
+ calculated according to the binary. It also implies the ``O`` flag.
+ This feature should be used with care as the interpreter
+ will run with root permissions when a setuid binary owned by root
+ is run with binfmt_misc.
+ ``F`` - fix binary
+ The usual behaviour of binfmt_misc is to spawn the
+ binary lazily when the misc format file is invoked. However,
+ this doesn't work very well in the face of mount namespaces and
+ changeroots, so the ``F`` mode opens the binary as soon as the
+ emulation is installed and uses the opened image to spawn the
+ emulator, meaning it is always available once installed,
+ regardless of how the environment changes.
+ ``T`` - transparent
+ Run the interpreter transparently. The binary is handed to
+ the interpreter through ``AT_EXECFD`` (``T`` implies ``O``),
+ the argument vector is left exactly as the caller built it
+ and the kernel labels ``/proc/pid/exe`` with the binary
+ instead of the interpreter. The interpreter has to load the
+ binary from ``AT_EXECFD`` and follow the
+ ``AT_FLAGS_TRANSPARENT_INTERP`` contract. Combining ``T``
+ with ``P`` is rejected: transparency preserves the whole
+ argument vector, argv[0] included.
+ ``L`` - loader substitution
+ Do not run the interpreter on the binary at all: load the
+ binary itself as a fully native exec and substitute the
+ interpreter for the loader named in the binary's
+ ``PT_INTERP``. See the "Loader substitution" section
+ below. ``L`` rejects ``T``, ``P``, ``O`` and ``C``;
+ ``F`` composes.
+ ``D`` - registered disabled
+ The entry is created disabled instead of being matchable at
+ once, and has to be enabled by writing ``1`` to its file
+ before it dispatches anything. This splits a registration
+ into creating the entry and activating it, leaving room to
+ configure it in between - which is what a ``B`` entry that
+ binds interpreters needs; see the bpf section below. The flag
+ is spent on the registration and is not read back: what an
+ entry file reports afterwards is whether it is enabled.
+
+
+There are some restrictions:
+
+ - the whole register string may not exceed 1920 characters
+ - the magic must reside in the first 128 bytes of the file, i.e.
+ offset+size(magic) has to be less than 128
+ - the interpreter string may not exceed 127 characters
+ - an interpreter used with ``C`` or ``L`` but without ``F`` has to be
+ named by an absolute path. It is opened when the binary is executed, so
+ a relative one would be resolved against the working directory of
+ whoever runs the binary
+ - the amount of pre-opened interpreters by ``F``, or bound to a ``B`` entry
+ is limited by the ``/proc/sys/user/max_binfmt_misc_interpreters`` sysctl. A
+ registration past the limit is refused with ``-ENOSPC``. This limits an
+ unprivileged namespace pinning files. A nested namespace can raise only its
+ own limit and every ancestor is charged too
+
+
+To use binfmt_misc you have to mount it first. You can mount it with
+``mount -t binfmt_misc none /proc/sys/fs/binfmt_misc`` command, or you can add
+a line ``none /proc/sys/fs/binfmt_misc binfmt_misc defaults 0 0`` to your
+``/etc/fstab`` so it auto mounts on boot.
+
+You may want to add the binary formats in one of your ``/etc/rc`` scripts during
+boot-up. Read the manual of your init program to figure out how to do this
+right.
+
+Think about the order of adding entries! Later added entries are matched first!
+
+
+A few examples (assumed you are in ``/proc/sys/fs/binfmt_misc``):
+
+- enable support for em86 (like binfmt_em86, for Alpha AXP only)::
+
+ echo ':i386:M::\x7fELF\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x03:\xff\xff\xff\xff\xff\xfe\xfe\xff\xff\xff\xff\xff\xff\xff\xff\xff\xfb\xff\xff:/bin/em86:' > register
+ echo ':i486:M::\x7fELF\x01\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x02\x00\x06:\xff\xff\xff\xff\xff\xfe\xfe\xff\xff\xff\xff\xff\xff\xff\xff\xff\xfb\xff\xff:/bin/em86:' > register
+
+- enable support for packed DOS applications (pre-configured dosemu hdimages)::
+
+ echo ':DEXE:M::\x0eDEX::/usr/bin/dosexec:' > register
+
+- enable support for Windows executables using wine::
+
+ echo ':DOSWin:M::MZ::/usr/local/bin/wine:' > register
+
+For java support see Documentation/admin-guide/java.rst
+
+
+You can enable/disable binfmt_misc or one binary type by echoing 0 (to disable)
+or 1 (to enable) to ``/proc/sys/fs/binfmt_misc/status`` or
+``/proc/.../the_name``.
+Catting the file tells you the current status of ``binfmt_misc/the_entry``.
+
+You can remove one entry or all entries by echoing -1 to ``/proc/.../the_name``
+or ``/proc/sys/fs/binfmt_misc/status``. A single entry can also be removed
+by simply unlinking (``rm``) ``/proc/.../the_name``.
+
+
+bpf-backed handlers
+-------------------
+
+With ``CONFIG_BINFMT_MISC_BPF`` both the matching and the interpreter
+selection can be delegated to bpf programs. A handler is an instance of the
+``binfmt_misc_ops`` struct_ops with a ``match`` and a ``load`` program and a
+``name``. Once the struct_ops map is registered the handler can be activated
+with a ``B`` entry that references it by name in the ``interpreter`` field
+and carries neither offset, magic, nor mask::
+
+ echo ':qemu:B::::my_handler:' > register
+
+Both programs receive the ``linux_binprm`` of the binary and both can
+sleep. The ``match`` program decides whether the handler applies: it is
+consulted during the entry walk exactly like magic and extension matching,
+in the same registration order with the same first-match-wins semantics.
+Unlike static matching it is not limited to the prefetched first bytes of
+the file in ``bprm->buf``: it can read the file, e.g. to parse ELF program
+headers whose data sits at arbitrary offsets. It only decides, though: the
+selection kfuncs below are rejected in it. The ``load`` program of the
+matched handler then selects the interpreter: it can equally read the file
+and derive the interpreter from the binary's location. It selects the
+interpreter by calling the ``bpf_binprm_set_interp()`` kfunc with an
+absolute path and returning ``0``. A match is committed: a failing
+``load`` fails the exec with its error instead of falling through to later
+entries; ``-ENOEXEC`` lets the remaining binary formats have a go. A path
+selected this way is opened with the credentials of the task doing the
+exec, exactly as a statically registered interpreter without ``F`` would
+be.
+
+An entry can instead bind the interpreters its handler may use, so that no
+path is resolved at exec time at all. An entry registered with ``D`` is not
+matchable yet, which is what leaves it open to being given them, one
+``+name path`` write at a time::
+
+ echo ':qemu:B::::my_handler:D' > register
+ echo '+aarch64 /usr/bin/qemu-aarch64' > qemu
+ echo '+arm /usr/bin/qemu-arm' > qemu
+ echo 1 > qemu
+
+Each path is opened during its write, in the writing process's context and
+with the credentials the entry file was opened with, exactly the way ``F``
+pre-opens a static entry's interpreter; the paths must be absolute. The
+path is everything past the first space, so there is nothing it cannot
+express, and no interpreter has to fit in a register string. An entry
+binds at most 100 interpreters, and each one is charged against
+``max_binfmt_misc_interpreters`` like any other binding. A write past either
+limit is refused with ``-ENOSPC``.
+
+The ``load`` program then selects one per exec by name with the
+``bpf_binprm_select_interp()`` kfunc, and every exec runs a clone of the
+file that was opened. The path decides which file is bound and nothing
+else: it is not resolved again, in any namespace, so what it holds later -
+or what it holds in the namespace of whoever runs the binary - no longer
+decides anything.
+
+Enabling the entry ends this. Its interpreters are read at exec time with
+nothing but a reference held on the entry, so an entry that has ever been
+matchable can never have its set changed again: the first ``1`` seals it,
+from then on ``+`` is refused with ``-EBUSY``, and an entry registered
+without ``D`` is sealed from the start. Binding a name twice is refused
+with ``-EEXIST``.
+
+Selection is by name so that the configuration and the program need not
+agree on an order, and so that a handler is not tied to where a distribution
+puts its interpreters. A name is a single word of printable ASCII, at most
+32 characters; a name the entry did not bind gives the program ``-ENOENT``,
+which it can act on or return. The interpreter runs under the path it was
+registered under, and the entry reports what it bound::
+
+ $ cat /proc/sys/fs/binfmt_misc/qemu
+ enabled
+ bpf my_handler
+ bpf-interpreter aarch64 /usr/bin/qemu-aarch64
+ bpf-interpreter arm /usr/bin/qemu-arm
+ flags:
+
+The path reported is the one the interpreter was bound under, which named
+the file at that moment; it is not re-resolved, so it is a record of what
+was bound rather than a promise about what that path holds now.
+
+The ``load`` program can also pass a single argument to the interpreter with
+the ``bpf_binprm_set_interp_arg()`` kfunc. It is inserted between the
+interpreter and the binary, exactly like the optional argument of a ``#!``
+interpreter line, e.g. for a handler that resolves ``$ORIGIN`` in a script's
+``#!`` path and needs to preserve the argument that followed it.
+
+The invocation flags a static entry fixes at registration - ``P``, ``C``,
+``O``, ``T`` and ``L`` - are per-exec choices for a bpf handler, made by the
+``load`` program with the ``bpf_binprm_set_flags()`` kfunc, so a single
+handler can decide them differently for each binary it handles:
+
+- ``BPF_BINPRM_PRESERVE_ARGV0`` keeps the caller's ``argv[0]`` (the ``P``
+ flag).
+- ``BPF_BINPRM_CREDENTIALS`` computes credentials from the binary (the ``C``
+ flag), bounded to user namespaces that map the binary's owner just like
+ any other setuid exec.
+- ``BPF_BINPRM_EXECFD`` opens the binary on the interpreter's behalf and
+ passes it through the ``AT_EXECFD`` aux vector entry (the ``O`` flag), so
+ the interpreter can run binaries it could not open by path.
+- ``BPF_BINPRM_TRANSPARENT`` runs the interpreter transparently (the ``T``
+ flag): the binary is handed over through ``AT_EXECFD`` as
+ with ``BPF_BINPRM_EXECFD``, but the argument vector is also left as the
+ caller passed it. An interpreter that loads the binary from ``AT_EXECFD``
+ then appears in ``argv[0]`` and ``/proc/pid/cmdline`` as a direct
+ execution of the binary. ``BPF_BINPRM_PRESERVE_ARGV0`` and a staged
+ interpreter argument are rejected in combination with it, just as ``P``
+ is with ``T``. It also lets a handler
+ run a binary passed as an inaccessible ``O_CLOEXEC`` file descriptor to
+ ``execveat()``, which a path-splicing dispatch cannot: the interpreter
+ has no path by which to open it.
+- ``BPF_BINPRM_LOADER`` substitutes the interpreter for the binary's
+ ``PT_INTERP`` and runs the binary as a fully native exec (the ``L``
+ flag). It excludes the other flags and a staged interpreter argument.
+
+Because these are program choices, a ``B`` entry carries no invocation
+flags in the register string; ``F`` has none to spell for it either, since
+the interpreters it binds already pre-open what ``F`` would. The
+registration directive ``D`` is the exception: it decides how the entry
+starts out, not how the interpreter is invoked.
+
+A handler is looked up only in the user namespace the struct_ops map was
+registered in. Handlers are not inherited, so an entry can only reference a
+handler registered in the same user namespace as its binfmt_misc instance.
+The entry keeps the handler alive; deleting the struct_ops map only prevents
+new activations.
+
+
+Transparent interpreters
+------------------------
+
+With the ``T`` flag or ``BPF_BINPRM_TRANSPARENT`` the dispatch is invisible
+to the resulting process. The argument vector is left exactly as the caller
+built it. The binary is passed through ``AT_EXECFD``. The kernel also labels
+``/proc/pid/exe`` correctly. The binary's file is write-denied while the
+process runs and the interpreter's is not, exactly as if the binary had been
+executed directly. A transparent entry does not change how credentials are
+derived. As
+with any other entry, set*id bits of the binary are only honored with ``C`` (or
+``BPF_BINPRM_CREDENTIALS``).
+
+The interpreter has to be built for this contract. The kernel announces it
+with ``AT_FLAGS_TRANSPARENT_INTERP`` in the ``AT_FLAGS`` aux vector entry
+next to ``AT_EXECFD``. The argument vector belongs entirely to the program,
+nothing was spliced in, so the interpreter doesn't consume arguments and
+simply loads the program from the descriptor. The bit is also the loader's
+license to finish the identity. After mapping the program it may retarget the
+``AT_PHDR``/``AT_ENTRY``/``AT_BASE`` entries of ``/proc/pid/auxv`` and the
+code/data statistics markers via one ``PR_SET_MM_MAP`` which completes
+what attaching debuggers observe. What remains visibly different from a
+direct execution is the address space layout. The interpreter occupies
+the main-image position and the program lives in the mmap region.
+
+
+Loader substitution
+-------------------
+
+The ``L`` flag turns the execution model around. Instead of running the
+registered interpreter with the binary as its payload the kernel loads
+the matched binary itself as the main image and substitutes the registered
+interpreter for the loader named in the binary's ``PT_INTERP``.
+
+Because the exec is native, there is no dispatch identity to
+reconstruct and no contract the substitute has to implement. A stock
+dynamic loader works unchanged. The argument vector is untouched,
+credentials and ``AT_SECURE`` derive from the binary, there is no
+``AT_EXECFD`` and no marker in the aux vector, the binary sits in the
+main-image slot with the native brk placement so ``/proc/pid/maps``,
+core dumps and perf mmap records have the native shape, and the
+identity is already complete when ``PTRACE_EVENT_EXEC`` stops the
+tracee. So launching under a debugger works, not just attaching. ``L``
+entries are for ELF binaries of a native architecture. Foreign-arch
+emulation and non-ELF payloads remain the domain of the classic and
+transparent modes.
+
+The override applies when the format that finally claims the file is
+ELF with a ``PT_INTERP``. A matched binary without one or an
+interpreter-less ``ET_DYN`` drops the override and runs natively. A file
+claimed by another format - a ``#!`` script, say - is handled by that
+format as if the entry had not matched. ``L`` is therefore not an
+enforcement mechanism: it decides how a binary that asks for a loader is
+run, it does not guarantee that everything matching the entry runs under
+the substitute. A format that cannot consume the override at all instead
+refuses the exec with ``ENOEXEC`` before the point of no return.
+
+A wrong-architecture ELF fails the whole exec with ``ENOEXEC`` exactly
+as if no entry had matched. A substitute that is not ELF of the right
+architecture fails with ``ELIBBAD``. The usual ``PT_INTERP`` sanity
+checks on the binary still apply. But the segment's content is otherwise
+irrelevant.
+
+``L`` rejects the classic-dispatch flags ``T``, ``P``, ``O`` and ``C``
+at registration. ``F`` composes and is valuable: with it the substitute
+is opened at registration time, so later mount namespace or path changes
+cannot redirect it. Without it the substitute is opened when the binary
+is executed, and the path is resolved in the mount namespace and root of
+whoever runs the binary, which is why it has to be absolute. As with
+``C``, register only trusted interpreters. The substituted loader runs
+with credentials derived from the binary.
+
+
+Hints
+-----
+
+If you want to pass special arguments to your interpreter, you can
+write a wrapper script for it.
+See :doc:`Documentation/admin-guide/java.rst <./java>` for an example.
+
+Your interpreter should NOT look in the PATH for the filename; the kernel
+passes it the full filename (or the file descriptor) to use. Using ``$PATH`` can
+cause unexpected behaviour and can be a security hazard.
+
+
+Richard Günther <rguenth@tat.physik.uni-tuebingen.de>