Most important differences from the old version:
- strip global process information from the thread struct into a
dedicated one,
- a parent is informed about the child death when the whole process
(the last thread) dies (not on each thread exit),
- all threads have the same parent (spawning thread is NOT the parent of
the spawned thread),
- a thread is able to wait on children created by another thread,
- a process is able to wait for exited children after execve,
- rewritten `waitid` implementation (no more gotos, supports __WCLONE
and friends flags),
- added option for syscall restarting, for now used only in `waitid`.
Additionally various bugfixes, cleanups and missing locks added.
Sometimes we need to temporarily stop IPC helper thread from receiving
more messages, e.g. when doing execve just before migrating exited (but
not yet waited for) children list.
* Make sure "stat.h" and "perm.h" are directly included where
necessary.
* Don't include "perm.h" inside "stat.h" but require it to be
included separately.
* Remove workarounds with __KERNEL__, __GLIBC__, defining pid_t
directly, and reversed include order (system headers before local
ones).
Logging to file was broken, because the PAL file write operation
required the user to provide an absolute offset, and LibOS always
provided an offset of 0. This worked when logging to stdout, but
in case of a regular file, it kept overwriting the beginning of
file.
To fix that, we introduce a a special DkDebugLog call. This is a
better solution than tracking the file offset manually, because
the offset would need to be synchronized across different threads
and processes, and debug logs should be as simple as possible. At
the same time, we don't want PAL to provide a generic "append to
a file" mechanism, because it makes I/O less deterministic.
The manifest syntax stays exactly the same, including 0 and 1
integers to denote boolean values (this is done for ease of porting
and can be fixed in future commits). The only visible change is
surrounding strings in the manifest with quotes (requirement of
TOML). All manifests and Makefiles of our tests and example apps are
ported to the new TOML syntax. Documentation is updated.
The operation loops indefinitely on error. Instead, it should find
the first free FD, and then try allocating it.
In addition, the right error after exceeding the limit is EMFILE
(however, dup2() is still supposed to return EBADF if asking for
an out-of-range value, as checked by the dup201 LTP test).
They all lacked error checking and `wait_event` was completely broken:
it was reading from non-blocking pipe and treating EAGAIN as
successfully waited-for event.
Previously, IPC_PORT_SERVER meant "listening port", and the actual
communication ports had several types. Only two of these types were
used for differentiation during IPC broadcast (direct-child and
direct-parent types). All other types denoted who is the remote party
this port connects to, but this info is superfluous. So this commit
replaces all these types with a generic IPC_PORT_CONNECTION, and
renames IPC_PORT_SERVER to a more familiar IPC_PORT_LISTENING.
Previously, LibOS (shim) layer of Graphene didn't support XSAVE area.
The XSAVE area stores FP, XMM, YMM, ZMM, etc. registers and control
states and is handled via FXSAVE/XSAVE and FXRSTOR/XRSTOR x86-64
instructions. This commit is the first step towards adding full-
fledged XSAVE support to LibOS. It adds XSAVE related structs and
functions to LibOS code, and propagates XSAVE regs/states from
parent to child on thread creation via clone() (though not really
correctly). New test `fp_multithread` is added to LibOS regression.
Future commits will add correct XSAVE handling on syscall transitions
and arriving signals.
If a remote ipc port gets disconnected we assume that remote process
died unexpectedly and mark it as killed with SIGKILL, so it makes no
sense to also keep the exit code.
On Linux fork and vfork are just specific cases of clone. This commit
does small cleanup of clone code and deduplicates proces copying code by
always using clone.
Removed:
- `message_confirm` - not used anywhere (and probably won't ever be),
- `SYS_PRINTF` - this was just synonym of `debug` or `warn` (depending
on the context),
- all other functions that were used only by the two above.
I don't know any reason why would stating the file name we're in be
helpful for anything. Moreover, this information was incorrect in a few
cases (copy-paste bugs, probably).
Additionally, a few minor type/formatting fixes included.
When compiling against v5.8 (or later) kernel headers, we encounter
compilation errors such as:
In file included from /usr/include/asm/unistd.h:13,
from shim_parser.c:11:
shim_parser.c:408:10: error: array index in initializer exceeds array bounds
[__NR_io_uring_setup] = {.slow = 0, .parser = {NULL}},
The remedy is to extend `syscall_parser_table` by modyfing the
definition of LIBOS_SYSCALL_BOUND so the offending system calls fit in
the table.
While at it, also fix the actual values for syscall numbers in
"shim_syscalls.h" (the numbers are not sequential, there is an
intentional gap between the numers).
Previously, Graphene preallocated 64MB for PAL internal metadata
like trusted/protected files metadata, handles metadata, etc.
If this limit was depleted, Graphene loudly failed, and the user
had no option but to change constant in source code and rebuild
Graphene. This commit adds the manifest option
`loader.pal_internal_mem_size` to allow increasing this limit.
This commit removes some unused code from LibOS, the biggest part of
which being the support for "inline" binaries linked directly against
LibOS.
This allows us to remove e.g. POINTER_TYPE() macro, which was supposed
to check if a type is a pointer type, but implemented the check based
only on the first 2 characters of the type name (?!?).
Instead, require the user to always provide a big enough buffer.
This is because inlined alloca() doesn't work in clang, see
https://github.com/oscarlab/graphene/issues/1794.
- Create a convenience "into qstr" wrapper
- Eliminate the only instances where the function was called with
on_stack == false (they were leaking memory)
- Remove a few null checks that are impossible now
- Fix possible overflow in unix_copy_addr() (truncate the address
instead)
Previously memory content was sent inside a checkpoint and counted
towards the checkpoint size in the child process, but not in the parent
process. The child allocated the checkpoint at the same address as in
the parent, but it had a different size, so it could colide with another
memory mapping. This commmit fixes it by making checkpointing code
receive the memory directly to destination addresses (which additionally
saves an unnecessary copy).
This commit contains the following changes:
- Removing macros like CONCAT, NS_SEND, NS_CALLBACK and replacing
with actual function/symbol names.
- Removing distinction between PID and SYS-V namespaces and their
corresponding leaders: all range requests for both namespaces
now go through the same logic and are served by the same set of
functions like ipc_lease_send/ipc_lease_callback, etc.
- Removing unnecessary IPC header files, by consolidating all
declarations in shim_ipc.h.
- Moving logic from shim_ipc_ns.h and shim_ipc_nsimpl.h to a new
shim_ipc_ranges.c -- responsible for alloc/dealloc of ranges.
In particular, this commit:
- Removes SLAB_DEBUG macros and corresponding code.
- Fixes memory leak in memmgr's enlarge_mem_mgr() by removing
__set_free_mem_area() call.
- Fixes bug of double-free of the very first memmgr area in
destroy_mem_mgr().
- De-duplicates "get new memory object" code by changing
get_mem_obj_from_mgr() to call get_mem_obj_from_mgr_enlarge().
- Simplifies and improves performance of free_mem_obj_to_mgr() since
there is no need to double-check that the object belongs to one of
the memmgr's areas because we already check memory_migrated().
- Fixes bug of free of wrong object in slabmgr's destroy_slab_mgr().