This commit improves the emulation of polling mechanisms (select,
pselect, poll, ppoll, epoll_wait) and cleans up the corresponding
code:
- New DkObjectsWaitEvents() PAL interface, replaces the inefficient
DkObjectsWaitAny() interface. This interface closely resembles
Linux/POSIX poll() in semantics.
- Improved shim_do_epoll_wait() implementation, now using the new
DkObjectsWaitEvents() interface.
- Improved shim_do_poll() implementation, now using the new
DkObjectsWaitEvents() interface.
- Small cleanups of polling code.
Accurate cleanup of shim_do_epoll_create1(), shim_do_epoll_ctl(),
shim_do_epoll_wait(), and other epoll helper functions. This cleanup
also adds error handling (missing previously).
The commit makes (most of) the corresponding LTP tests pass now.
Note that epoll semantics are still incorrect and inefficient: current
epoll_wait() emulation returns only one event to the user.
Graphene IPC (GIPC) was introduced to perform faster bulk IPC by sharing pages as copy-on-write
across processes. This feature became stale, and it was shown that recent Linux kernels (4.2+)
have zero-copy transfers over UNIX sockets and exhibit similar performance. GIPC is not built and
not tested in Jenkins. Also, GIPC does not work under SGX. This commit completely removes GIPC.
This commit completely rewrites futex implementation to (hopefully)
remove all races (both on memory access level and between waits and
wakes), make it compatible with actual Linux implementation and make
it more maintainable and readable.
Previously, if user performed getdents() on a directory containing
inaccessible files (because user doesn't have permission), whole
getdents failed with -EACCES. This is incorrect behavior: files must
still be listed. This commit fixes the root cause of this bug by
marking inaccessible files as DENTRY_NEGATIVE.
Previously, there was a data race on thread::is_alive between one thread
checking whether it is the last thread alive via check_last_thread() and
another thread exiting via thread_exit(). The former checks if is_alive
is true, the latter sets it to false. However, the exiting thread will
truly exit only after it called DkThreadExit(), thus the race on is_alive
led to scenarios where two threads believe to be the last threads alive
and compete on terminating Async Helper/IPC threads and exiting the whole
process. This commit introduces cleanup_thread() called by Async Helper
to set is_alive to false and delete the thread, freeing its resources.
The data race is thus removed, and shim_thread object leak is prevented.
When child thread exits, it wakes up its parent if CLONE_CHILD_CLEARTID was set
during clone() call. Previously, this was done by the child thread itself as
part of its own clean-up in release_clear_child_id(). But this child thread is
still alive at this point and uses some resources, most notably the stack (that
might have been provided by the parent) and the SGX TCS slot. Upon waking up,
the parent might decide to free that stack (as Pthreads do) or re-use the TCS
slot, causing data races.
This commit introduces a correct emulation of CLONE_CHILD_CLEARTID:
- A new argument `PAL_PTR clear_child_tid` is added to DkThreadExit();
it points to internal Graphene memory that is erased on child exit to notify
Async Helper thread.
- At PAL layer, when thread finally exits, it sets PAL-level *clear_child_tid = 0
(corresponds to &clear_child_tid_val_pal at LibOS level); this signals to LibOS
layer that the thread stopped using resources.
- At LibOS layer, Async Helper thread is set up to wait for the signal from
PAL; it is now the responsibility of Async Helper thread to call
release_clear_child_id() to wake up the parent thread.
- Async Helper thread waits for clear_child_tid_val_pal == 0 and then sets
the actual clear_child_tid to 0 and wakes up the waiting parent.
Note that for Linux-SGX PAL, clear_child_tid is set to 0 not immediately
but as part of handle_thread_reset, otherwise the TCS slot could be still
occupied when LibOS wakes up the parent.
As a side effect, the LibOS code for threads/process exit is cleaned up.
This commit also fixes all regression tests to use the new signature of
DkThreadExit() and increases the number of SGX threads slightly (to
accommodate the newly used Async Helper thread).
This patch removes __attribute__((packed)) to eliminate warnings gcc-9
generates. Example:
> warning: taking address of packed member of struct poll_handle may result in an unaligned pointer value [-Waddress-of-packed-member]
Packed attribute is used for structures using which Graphene processes
communicate with each other. We don't implement privilege separation, so
it's ok to leak data in struct paddings.
If the padding is really a concern, we could:
- Define two structures, a packed one for serialization and a non-packed
one for code, and then explicitly (de)serialize.
- Define the structure fields with an explicit size (e.g. uint64_t
instead of long), carefully add explicit padding and then zero it out
before sending.
The execve() syscall starts a new executable in the *same* process. Previously,
Graphene followed this convention *only* for non-SGX PALs. If PAL was Linux-SGX,
Graphene silently terminated the process and created a new one.
This deviation from standard execve() behavior resulted in the host shell
becoming detached from the Graphene-SGX process. In turn, this led to our
Bash example (on Ubuntu 18.04, bash version 4.4.19) "terminating" early
from the point of view of Jenkins, and SGX-18.04 pipeline failed.
This commit allows Graphene-SGX to execve() in the same process, but only if it
is the same executable (as in the Bash example). It is still impossible to
execve() in the same process for a different executable since this requires a
new SGX enclave measurement and thus demands a new process.
Now, shim_thread::tcb is used only as an integer value for %fs_base,
and shim_thread::user_tcb is not needed anymore. This commit renames
shim_thread::tcb to shim_thread::fs_base and removes user_tcb.
For binaries statically linked against Glibc, %fs register cannot be
used for LibOS TCB (shim_tcb) because it is already used by Glibc.
Thus, this commit moves shim_tcb into PAL TCB. Also, now that LibOS
doesn't access Glibc TCB (__libc_tcb), we make it an opaque pointer
used only for clean up.
This commit adds dummy implementations for setpriority, getpriority,
sched_setparam, sched_getparam, sched_setscheduler, sched_getscheduler,
sched_get_priority_max, sched_get_priority_min, sched_rr_get_interval,
sched_setaffinity, sched_getaffinity. These implementations only check
for incorrect user-supplied arguments and either do nothing (setters)
or return default values (getters). This commit also adds a simple LibOS
regression test.
This commit adds support for eventfd():
- shim_do_eventfd() and shim_do_eventfd2() emulation at LibOS level;
- new eventfd pseudo-FS;
- db_eventfd.c emulation to route eventfd calls to host OS at Pal
level (implementation for Linux and Linux-SGX, stubs for Skeleton);
- new OCALL ocall_eventfd() for Linux-SGX Pal;
- LibOS regression test for eventfd.
This implementation of eventfd() correctly handles IPC between
threads of the same process; it also must handle IPC between parent/
child processes. This implementation currently doesn't support the
scenario when kernel signals the process via eventfd, but is easily
extensible for this. The implementation currently doesn't have
additional checks to prevent Iago attacks on eventfd.
There were some places where put_thread() calls were missing. This
caused reference counter to never reach 0, so that the shim_thread
struct was never freed, thus leaking memory.
- Make assert() a no-op in non-debug builds.
- Use static_assert for compile-time asserts.
- Fix assert() implementation (previous version didn't work for
expressions with types larger than long, it also always printed
`(value:0)`).
- Clean up calls to asserts.
In test_user_memory(), a memory range is tested via probing of each page
in the range. If the memory page was not allocated, it leads to a
segfault which is captured by the LibOS memfault_upcall() and reported
in the variable tcb.test_range.has_fault. However, the compiler may
optimize away accesses to this variable in test_user_memory(), believing
it is never updated anywhere else. This commit introduces a memory
barrier to prevent this compiler optimization (same for test_user_string()).
Inside `struct shim_handle` `opened` reference counter is used incorrectly.
Additionally at this moment it guards the same resource (handle) as `ref_counter`,
making it obsolete. This patch removes `opened` counter and moves `close_handle`
logic into `put_handle`.
Internal LibOS and PAL interfaces use microseconds (us) for timeout
values. However, Linux epoll_wait/epoll_pwait syscalls use milliseconds
(ms) for timeout. Previously, there was a bug in timeout resolution
because epoll_wait() emulation did not convert from ms to us. This
commit fixes this bug and also adds suffixes "_ms" and "_us" to make the
time units used explicit.
enable_preempt(), disable_preempt(), and other functions operated on
shim_context.preempt using non-atomic operations. This commit replaces
the old broken implementation with atomic operations.
- Deprecate sys_stack_size and max_brk_size. Get the values directly from __rlim.cur.
- Add internal routines for setting and getting __rlim.cur.
- Implement prlimit64() and simplify getrlimit() and setrlimit().