Full text
D3.7: Solutions for Security Testing of Low-Level System Components (first version) PROJECT Project Number 101120962 Project Acronym RESCALE Project Title Revolutionised Enhanced Supply Chain Automation with Limited Threats Exposure Start Date 01.10.2023 Programme HORIZON-CL3-2022-CS-01-02 DELIVERABLE Deliverable Type R - Document Report Workpackage WP3 Deliverable Lead VUA Editors Cristiano Giuffrida Contributors ISI, INT, DIE, BNR Dissemination Level PU - Public Abstract This deliverable presents the first version of the RESCALE solutions that target security testing of low-level system components. Specifically, we outline the features, functionalities, and architecture of the core solutions developed within Task 3.4. Our solutions focus on both software and hardware vulnerabilities that affect low-level system components. On the software side, we introduce SafeFetch, a solution to detect and mitigate double-fetch vulnerabilities–common issues in low-level system components. We also introduce FATex, a solution to perform firmware analysis for the purpose of security testing. On the hardware side, we present InSpectre Gadget, a solution to identify and analyze the exploitability of Spectre-type vulnerabilities, which remain a significant and poorly understood attack surface for low-level system components. Disclaimer The information in this document is provided “as is”, and no guarantee or warranty is given that the information is fit for any particular purpose. The content of this document reflects only the author’s view – the European Commission is not responsible for any use that may be made of the information it contains. The users use the information at their sole risk and liability. This project has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No 101120962
D3.7: Solutions for Security Testing of Low-Level System Components (first version) Document Revision & Quality Assurance Internal Reviewers 1. Tzortzia Koutsouri - (CBRL) 2. Nassos Bountioukos - (CBRL) 3. Narges Yousefnezhad - (BNR) 4. Andrei Costin - (BNR) Revisions Version Date Partner Overview 0.1 20/07/2024 VUA ToC 0.2 08/11/2024 VUA First draft 0.3 13/11/2024 VUA, BNR Finished draft 0.4 19/11/2024 CBRL, BNR Reviewer comments 0.5 20/11/2024 VUA Addressed comments 0.7 03/12/2024 VUA Revised version 0.8 06/12/2024 VUA Added Interfaces 0.9 10/12/2024 ISI, AEGIS Final comments 1.0 13/12/2024 VUA Final version RESCALE – PU - Public – Page 2 / 50
Table of Contents 1 Introduction 8 1.1 Scope&Contribution............................... 8 1.2 Relation to Work Packages, Deliverables, and Activities . . . . . . . . . . . . . 8 1.3 Contribution to WP3 and Project Objectives . . . . . . . . . . . . . . . . . . . 8 1.4 DocumentStructure................................ 8 2 High-level Solution Overview 10 3 SafeFetch 12 3.1 Background .................................... 13 3.2 ThreatModel ................................... 14 3.3 Overview ..................................... 14 3.4 ProfilingKernelFetches.............................. 16 3.5 CacheFrontend .................................. 17 3.5.1 Efficient Range Queries . . . . . . . . . . . . . . . . . . . . . . . . . . 17 3.5.2 QueryResolution............................. 19 3.6 CacheBackend .................................. 19 3.6.1 CustomAllocator............................. 20 3.6.2 DualRegionDesign ........................... 21 3.6.3 Lifecycle Management . . . . . . . . . . . . . . . . . . . . . . . . . . 22 3.7 Interconnection/Interfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 3.8 RelatedWork ................................... 25 4 FATex 27 4.1 Architecture.................................... 27 4.2 Implementation .................................. 27 4.3 EvaluationResults................................. 28 4.4 Interconnection/Interfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 5 InSpectre Gadget 30 5.1 Background .................................... 31 5.1.1 Transient Execution Attacks . . . . . . . . . . . . . . . . . . . . . . . 31 5.1.2 Spectrev2................................. 31 5.1.3 Branch History Injection . . . . . . . . . . . . . . . . . . . . . . . . . 32 5.1.4 Defenses ................................. 32 5.2 ThreatModel ................................... 34 5.3 Overview ..................................... 34 5.4 InSpectreGadget ................................. 35 5.4.1 StandardGadgets............................. 35 5.4.2 Exploitation-Aware Gadget Analysis . . . . . . . . . . . . . . . . . . . 36 5.4.3 Design................................... 38 5.4.4 Limitations ................................ 40 5.5 Interconnection/Interfaces . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40 5.6 RelatedWork ................................... 41 6 Main Innovations & Conclusion 43 6.1 MainInnovations ................................. 43 3
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 6.2 Conclusion .................................... 43 6.3 FutureWork.................................... 44 RESCALE – PU - Public – Page 4 / 50
List of Figures 1 Workflow of a double-fetch exploit. . . . . . . . . . . . . . . . . . . . . . . . . . 14 2 SafeFetch hindering a double-fetch exploit. . . . . . . . . . . . . . . . . . . . . . 15 3 SafeFetch’s high-level architecture. . . . . . . . . . . . . . . . . . . . . . . . . . . 16 4 Percentage of syscall samples vs. number of ranges they fetch. . . . . . . . . . . . 18 5 Percentage of syscall samples vs. average size of data they fetch. . . . . . . . . . . 20 6 Percentage of syscall samples vs. total amount of data they fetch. . . . . . . . . . . 21 7 Percentage of processes vs. number of syscalls fetching user data. . . . . . . . . . 22 8 Heatmap showing the total amount of user data fetched by fetch-heavy syscalls. . . 23 9The BHI attack. The attacker first triggers the branch CAwith history HA1 , which inserts the target TAinto the BTB 2 , then triggers the victim branch CV with history HV3 . The histories are crafted so that CAand CVshare the same BTB entry, so from CVthe CPU speculates to TA4 .................. 32 10 InSpectre gadget workflow. The analyst provides a kernel image and a list of target addresses to InSpectre Gadget 1 , which performs in-depth inspection to find gadgets that can leak secrets and output their characteristics. The gadgets can be filtered 2 based on the available attacker-controlled registers and the mitigations enabled, and used to craft Spectre-v2 exploits against the kernel 3 . ........ 35 11 Known-prefix technique. The attacker first points the secret address to some known data 1 . Then, by shifting the address 2 , small portions of the secret are revealed. With that, one can adjust the base of subsequent transmissions 3 . . . . . 37 12 Anatomy of a transmission. Once a potential transmission is identified through symbolic execution, InSpectre Gadget dissects its symbolic expression into components, which are then analyzed to reason about exploitability. . . . . . . . . . . 40 5
D3.7: Solutions for Security Testing of Low-Level System Components (first version) List of Abbreviations AST Abstract Syntax Tree. 39 BHB Branch History Buffer. 32, 33 BHI Branch History Injection. 30–34 BTB Branch Target Buffer. 31, 32 BTI Branch Target Injection. 30, 31 CLI Command Line Interface. 28 CPU Central Processing Unit. 30–34, 38 CSV Comma Separated Values. 28 CSV2 Cache Speculation Variant 2. 31 CVE Common Vulnerabilities and Exposures. 31 DSCG Dynamic Supply Chain component Guarantee. 8, 10, 28, 43 eBPF Extended Berkeley Packet Filter. 30–34 eIBRS Enhanced Indirect Branch Restricted Speculation. 30, 31, 33, 34 FAT Firmware Analysis Toolkit. 27 FATex Firmware Analysis toolkit extended. 28, 29 FineIBT Fine-Grained Indirect Branch Tracking. 30, 34 HTML HyperText Markup Language. 28 IBT Indirect Branch Tracking. 30, 33, 34 IoT Internet of Things. 28 JSON JavaScript Object Notation. 28 MMU Virtual Memory Area. 12, 16 OS Operating System. 12 PTE Page Table Entry. 13 QEMU Quick Emulator. 27, 29 SLAM Spectre based on Leaky Address Masking. 38 RESCALE – PU - Public – Page 6 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) SMAP Supervisor Mode Access Prevention. 13 SMEP Supervisor Mode Execution Prevention. 13 SMT Simultaneous Multithreading. 38 TBOM Trusted Bill of Materials. 28 TCB Trusted Computing Base. 15 TLB Translation Lookaside Buffer. 38 TOCTTOU Time-Of-Check To Time-Of-Use. 12, 14 VMA Virtual Memory Area. 16, 17 XSS Cross-Site Scripting. 28 RESCALE – PU - Public – Page 7 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 1 Introduction To complement the core static and dynamic analysis solutions developed in Task 3.1 and Task 3.2, respectively, as well as the physical side-channel leakage analyzer developed in Task 3.3, Task 3.4 focuses on developing solutions for security testing of low-level system components. These solutions emphasize program analysis and instrumentation of low-level systems software, such as operating systems kernels, to find low-level vulnerabilities and analyze their impact. As part of the overall RESCALE platform, these solutions perform hardware/software attack surface analysis and generate reports contributing to the generation of the Dynamic Supply Chain Component Guarantee (DSCG) (whose foundation is described in D3.3). 1.1 Scope & Contribution This document outlines the outcomes of the first phase of Task 3.4, Low-Level System Security Testing. This phase encompasses the initial design of the solutions for security testing of low-level system components. Our focus is on the core solutions targeting both software and hardware vulnerabilities that affect low-level system components on commodity platforms. The document is based on two papers developed within the RESCALE project and presented at the 2024 USENIX Security Symposium [73,28]. 1.2 Relation to Work Packages, Deliverables, and Activities Deliverable D3.7 results from Task 3.4 within Work Package 3 (WP3) and is also relevant to Work Packages 4 (WP4) and 5 (WP5). This deliverable details the design of the solutions for low-level system security testing, including specific integration aspects related to WP5. 1.3 Contribution to WP3 and Project Objectives Deliverable D3.7, which documents our progress in Task 3.4, aligns with WP3 Objective O3.5 to provide a platform for performing security tests for low-level software and processor implementations. Additionally, it contributes to Project Objective O2 by analyzing microarchitectural attacks. 1.4 Document Structure This section outlines the structure of the deliverable as follows: •Section 2 provides a high-level overview of the solutions for security testing of low-level system components. RESCALE – PU - Public – Page 8 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) •Section 3 details the design of SafeFetch, our solution for detecting and mitigating double-fetch (software) vulnerabilities. These vulnerabilities frequently affect low-level system components, such as operating system kernels. •Section 4 presents FATex, our solution for analyzing firmware, with support for static and dynamic firmware analysis for the purpose of security testing. •Section 5 describes the design of InSpectre Gadget, our solution for detecting and analyzing Spectre-type (hardware) vulnerabilities. These vulnerabilities represent an increasingly large attack surface for low-level system components. •Section 6 summarizes and concludes the deliverable. RESCALE – PU - Public – Page 9 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) Cache Backend Cache Frontend SafeFetch Syscall Cache syscalltransfer function call search range range query miss Custom Allocator sanitized range provision Figure 3: SafeFetch’s high-level architecture. implementation on modern operating systems such as Linux, faring even better than the state of the art (Midas). Third, our fetch-side instrumentation strategy may end up copying more data than Midas’ write-side strategy in case user data is never changed during syscall execution. However, we found that such cost is marginal compared to that of MMU-based instrumentation, resulting in better performance. Finally, our design require instrumenting the kernel’s fast path (i.e., transfer functions). As such, its instrumentation and data structures need to be carefully designed to efficiently support typical kernel fetch patterns. Nonetheless, we found it is possible to capture a variety of different fetch patterns with relatively simple data structures. In the next sections, we first analyze the patterns relevant to our design. Then, we use the insights gathered from our analysis to detail our design. 3.4 Profiling Kernel Fetches While our syscall caches are superficially similar to other kernel caches since they may support similar range queries, our design requirements are fairly unique. For instance, the Virtual Memory Area (VMA) cache is a classic example of a kernel cache supporting range queries, however, its lookup patterns are wholly different from ours (lookups on locality-friendly memory management operations vs. lookups on kernel fetches) and so are its scope (process vs. syscall) and data storage requirements (fixedvs. variable-sized data). As such, to make optimal design decisions, we need to learn more about typical patterns for kernel fetches, including RESCALE – PU - Public – Page 16 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) their frequency, data transfer size, etc. To this end, we developed a simple profiler to gather kernel-fetch statistics. Specifically, our profiler interposes on syscall execution and records the following statistics: the total number of ranges a syscall transfers from user space, the average size of ranges transferred by a syscall, and the total amount of data a syscall fetches from user space. Moreover, for each process executed during profiling, we also gather the number of syscalls that transfer data from user space. We use various benchmarks (e.g., LMBench, OSBench) and popular user applications (e.g., Nginx, Apache) to generate a workload to sample syscall execution. In total, our workload generated around 317 million syscall samples, exercising 165 individual syscalls (≈ 52% of all defined syscalls for Linux x86 64). Our main profiling results are depicted in Figures 4, 5, 6, 7, 8. We elaborate on the results in the next sections, using the gathered insights to motivate our design. 3.5 Cache Frontend To provide the kernel with a consistent view of user memory, the cache frontend interposes on all user-to-kernel transfer function invocations requesting a specific user range. In response, the frontend queries the syscall cache for the range by means of the user (virtual) address and the length of the range. After the query completes, the frontend performs a query resolution step, fetching parts of the range from user space into the cache if needed (i.e., if not cached) and then forwarding a sanitized range to the original transfer function. The frontend considers a user range A sanitized if and only if: for any sub-range Bof contiguous user addresses, such that B⊆A, and Bwas previously fetched during the execution of the syscall (i.e., via a previous transfer function) then Bmust consist of the same bytes as when it was first fetched. When querying the cache, the frontend uses a predetermined search policy to locate the range in the cache. The search policy is subject to the (meta)data structure used to bookkeep the user ranges in the cache. 3.5.1 Efficient Range Queries Given a user start and end address (i.e., a range) SafeFetch needs to find all cached chunks that overlap with this input range. Since a range may only partially overlap with an existing range (or multiple cached ranges), we are interested in finding the optimal data structure that can service this operation efficiently. For this purpose, other kernel subsystems use either linked lists (for small caches) or red-black trees (for large caches). The Virtual Memory Area (VMA) cache is case in point, generally serving address range queries via a per-process red-black tree. However, the VMA cache’s fast path uses a linked list for the few recently used VMAs. We experimented with both types of data structures in the context of SafeFetch. In both cases, a node contains metadata recording the start/end address of the range and a reference to the cached data. Clearly, in the case of many cached ranges, we expect linked-list-based queries to perform poorly, with a worst-case search complexity of O(n). In the same vein, we expect redblack trees to be more efficient, with a worst-case search search complexity of O(log(n)) due to constant-time rebalancing. Indeed, we experimentally verified that after around 100 cached ranges, the average search time of a linked list is slowed down by a factor of two compared to a RESCALE – PU - Public – Page 17 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 0 fetches 1 fetch [2-10) fetches [10-30) fetches [30-4099) fetches 0% 50% % of syscall samples 63.93% 28.69% 7.29% 0.03% 0.06% Figure 4: Percentage of syscall samples vs. number of ranges they fetch. red-black tree. However, when the cache contains only a few elements we observed the linked list significantly outperforming a red-black tree. On top of more lightweight search logic, a small linked list has another important performance edge over a small red-black tree on the insertion path (i.e., when adding a new user range into the cache). Indeed, due to rebalancing, insertions in the red-black tree are always around 10 orders of magnitude slower than in a linked list. Since according to our profiling results, double fetches are rare (occurring once every 1,277 sampled syscalls), insertions are frequent and are thus important to consider for performance. Selecting the ideal data structure To select the ideal candidate between our two data structures, we turn to our profiling results. Figure 4 shows the percentage of syscall samples that fetch Nranges from user memory across all the profiled benchmarks. As shown in the figure, most samples (≈64%) do not fetch user ranges at all, while the vast majority of those which do, fetch at most one range (≈28.7% of samples). Even so, a nontrivial number of samples (≈7.32%) fetch at most 30 ranges, while only ≈0.06% samples fetch over 30 ranges. As our results suggest, a (small) linked list seems the ideal candidate to support range queries for the vast majority of syscall samples. However, we also observed samples fetching as many as ≈4,000 ranges. In those cases, the linked list has very poor performance and the red-black tree is a vastly better option. To optimize for all possible syscall scenarios, we ultimately opted for an adaptive search policy. In other words, the search policy uses a linked list by default until the number of ranges in the cache reaches a predetermined threshold. When the threshold is exceeded, SafeFetch switches to a red-black tree implementation. To convert the linked list into a red-black tree, SafeFetch uses an efficient algorithm that constructs a balanced red-black tree from an ordered linked list. This is done by first copying the pointers to ranges in a vector and then iterating over the vector in a binary breath-first search fashion. To keep the algorithm efficient, we need to initiate the conversion when the list contains a number of ranges of the form 2N−1. RESCALE – PU - Public – Page 18 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 3.5.2 Query Resolution When processing the result of the query, the Cache Frontend makes different decisions depending on whether it found a cached user range colliding with the queried range (cache hit) or not (cache miss). In case of a miss, no subrange from the queried range was previously fetched from user space. As a resolution, the frontend fetches the range from user space and instructs the backend to allocate storage to insert the range into the cache. It then also indexes the new range by inserting new metadata into the linked list or red-black tree with with O(1) complexity—by piggybacking on the result of the previous query. Cache hits can either be perfect or partial. A perfect hit means that the frontend found a cached range which contains the entire range it queried for. In this case, the frontend forwards the sane range from the cache without performing any fetch from user space. A cache hit is partial if the frontend finds a cached range that collides with the range it queried for, but does not contain the entire range. In this case, the queried range may collide with multiple cached ranges previously fetched from user space and the frontend needs to first execute a range defragmentation step to determine the sanitized range. Let Bbe the queried range and let M={Ai|Ai∈cache∧AiTB= /0}containing all cached ranges that collide with B. Defragmentation involves computing a new range Csuch that C=∪n i=1Ai∪(B\ ∪n i=1Ai). In other words, the defragmented range C contains all bytes from the cached ranges colliding with B, while the sub-ranges of bytes that are not cached yet are fetched from user space. The frontend instructs the backend to replace all colliding ranges in the cache with the newly defragmented range C, after which it can service the sanitized range from C. Defragmentation piggybacks on the result of the previous search, because all colliding ranges can be found by using the previous search’s iterator. To support this, insertion in the linked list and red-black tree preserves the virtual address ordering of cached ranges. 3.6 Cache Backend The Cache Backend is responsible with managing the backing memory for the syscall cache. Specifically, its main goals are to efficiently manage the lifecycle of ranges in the cache and exploit CPU cache locality as much as possible. The latter can be achieved by stacking together ranges in memory for data locality and thus better CPU cache utilization, which speeds up range queries. The former involves enforcing policies for range allocation/deallocation and storage provisioning/relinquishing techniques to keep cache operations optimal. To achieve its goals, the backend maintains one cache for each in-transit syscall (at the perthread granularity) and adheres to a cache organization specifically tailored to maximize data locality when performing range queries through the cache. Additionally, the backend enforces a series of range lifecycle management policies to efficiently oversee cache lifespan. To support this overall strategy, the backend relies on a custom memory allocator. RESCALE – PU - Public – Page 19 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 2022242628210 212 214 Average amount of bytes per fetch 0.0% 5.0% 10.0% 15.0% 20.0% 25.0% 30.0% 35.0% 40.0% % of syscall samples Figure 5: Percentage of syscall samples vs. average size of data they fetch. 3.6.1 Custom Allocator A naive allocation policy would be to create an object in the cache every time a syscall fetches a new range, by using an of-the-shelf kernel allocator (e.g., slab) to accommodate range data and metadata. However, standard kernel allocators do not give control over where objects are placed in (virtual) memory, so we cannot assure locality to improve the performance of range queries. Moreover, for an off-the-shelf allocator the allocation and deallocation logic might cause non-trivial overhead when syscalls transfer many ranges from user space, which happens in practice (see Figure 4). A better approach is to service kernel memory in larger chunks to fit multiple ranges, i.e., using a region-based allocation scheme [4]. In such a scheme, a region consists of one or more buffers (contiguous memory blocks) with the same data lifetime. In a region, memory objects are allocated one after another in memory (in the current buffer) and get deallocated all at once (by flushing all the buffers), once the lifetime of all the objects in the region ends. As such, region-based allocation allows us to pack ranges together in memory to improve locality. Moreover, it reduces the number of calls to the underlying allocator when allocating ranges. On the fast path, an allocation involves only bumping a pointer to the next slot in the current buffer. Finally, it can efficiently deallocate all the allocated objects in one blow. SafeFetch uses a custom region-based allocator, which services kernel memory at the granularity of a region, every time the backend requests more storage to hold ranges. To understand how well SafeFetch can benefit from region allocation’s locality-friendly design, we turn again to our profiling results. Figure 5 shows the percentage of syscall samples that fetch an average size Sof user data across all the profiled benchmarks. As shown in the figure, the majority (i.e., 65%) of syscall samples fetch ranges that are on average less than 64 bytes, confirming a high degree of data locality in a region in the common case. SafeFetch also benefits from the fast deallocation path of region-based allocation, efficiently deallocating all the ranges in the region when the syscall terminates. As an optimization, SafeFetch does not discard the buffers once the region is deallocated, but adds them to a pool for (fast) reuse in new regions created RESCALE – PU - Public – Page 20 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 20232629212 215 Total amount of bytes copied from user. 0.0% 5.0% 10.0% 15.0% 20.0% 25.0% % of syscall samples Figure 6: Percentage of syscall samples vs. total amount of data they fetch. by future syscalls. The next question is the buffer size one should use. As Figure 5 suggests, small allocation requests are common suggesting small buffers are desirable. Moreover, Figure 6 shows the percentage of syscall samples that fetch a total amount of Bbytes of user data across all the profiled benchmarks. As shown in the figure, although a moderate portion of syscalls transfer much data (even above 8 pages), the vast majority transfer far less than a page of data. As a result, SafeFetch uses a single 1-page buffer by default in each new region (syscall) and elastically adds buffers as needed in the edge cases—many fetches per syscall or fetches transferring over 4 KB of data. 3.6.2 Dual Region Design A naive approach would be to store the user memory ranges and the metadata necessary for range queries as a standalone object in one single per-syscall region. However, this approach would lead to metadata fragmentation and poor locality because metadata would be intermixed with data bytes. While most fetched ranges are small, we saw that syscalls can transfer larger ranges as well (Figure 5). To maximize data locality for range querying, the Cache Backend partitions the syscall cache in two separate regions: a data region stores all byte ranges copied from user space while a metadata region stores the bear bone necessities to perform queries over ranges. Consequently, for each user range, the backend maintains two objects: a) a data object storing the range of bytes copied from user space and b) a metadata object storing the properties of the range used when querying (e.g., user virtual address, length, pointer to the data object, and a field used to link into a red-black tree or linked-list). The size of metadata objects is fixed and small, allowing one to provision metadata regions with buffers smaller than a page. Nonetheless, we chose to serve 1-page buffers to metadata regions as well because this leads to better CPU cache coloring (and utilization) [44]. Despite the same default buffer size, the two regions provision their buffers from separate memory RESCALE – PU - Public – Page 21 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 1 syscall [2-10) syscalls [10-50) syscalls 50 or more 0% 50% % of processes 0.4% 85.2% 5.6% 8.9% Figure 7: Percentage of processes vs. number of syscalls fetching user data. pools (i.e., slab caches) to further improve locality. 3.6.3 Lifecycle Management Range allocation policy A range is allocated in the cache every time the frontend encounters a cache miss while performing a range query. Our profiling data (e.g., Figure 4) suggests that most fetches are not double fetches, hence we expect frequent range allocations in the cache. Our custom allocator helps reduce the pressure on the (less efficient) underlying allocator, because many ranges can be allotted from the same region buffer. Allocating ranges entails creating metadata and data objects in the appropriate regions. To this end, the backend uses a region accounting structure, which keeps track of all buffers allotted to the referent region. To speedup object creation, the region accounting structure stores a pointer to a region head buffer and favors servicing allocations from this buffer. As region heads get depleted at some point, new buffers are added to the region when appropriate and region heads get updated. If an object cannot be created from the region head, the Cache Backend goes on the slow path iterating over all buffers allotted to the region until it finds one with enough space to service the request. As shown in Figure 6, some syscalls can transfer large amounts of user data, thus it is possible that a region can hold many depleted buffers. While this is more likely for data regions, it can also occur for metadata regions when syscalls execute many fetches. To optimize the slow path, the region accounting structure keeps a freelist containing only the buffers that are not yet depleted and can still service allocations. Lastly, if an object cannot be serviced on the slow path, then the backend leverages the custom allocator to create a new buffer in the referent region. RESCALE – PU - Public – Page 22 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 0 bytes [1,64) bytes [64, 256) bytes [256-1024) bytes [1024-4096) bytes [1-4) pages [4,16) pages above 16 pages write pwrite64 writev sendto execve 0.01 0.17 0.40 0.10 0.13 0.01 0.18 0.01 0.03 0.56 0.21 0.10 0.07 0.02 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00 1.00 0.00 0.80 0.00 0.00 0.00 0.00 0.20 0.00 0.00 0.00 0.00 0.00 0.57 0.39 0.03 0.00 0.0 0.2 0.4 0.6 0.8 ratio Figure 8: Heatmap showing the total amount of user data fetched by fetch-heavy syscalls. Cache invalidation policy When a system call terminates, all ranges currently held in the cache can be deallocated. As a result, one could relinquish all storage held by cached ranges on syscall exit. This strategy reduces the memory footprint and also yields a simple deallocation path flushing all the buffers held by metadata and data regions. Moreover, as discussed, syscalls do not fetch many ranges, thus on average we need not free many buffers. However, as shown in Figure 7, our profiling data shows that 99.6% of the processes that incur fetches do so across at least two syscalls. In other words, if a process issues a syscall that fetches user data, it is likely to issue other syscalls that do the same. Given our profiling results, it may be tempting to preserve all the allocated region buffers across syscalls. However, while this strategy minimizes the number of calls to the underlying allocator (improving performance), it may also significantly increase the steady-state memory footprint. That is particularly the case for threads issuing syscalls that fetch a lot of user data (even more than 8 pages, as shown in Figure 6). Given these observations, our cache invalidation policy is to preserve the head buffers, for both metadata and data regions, across the syscalls of a given thread but release all the other buffers held by each region, on syscall exit. Additionally, on each syscall exit, we invalidate the bookkeeping structure used by our frontend when performing range queries. Finally, on thread exit we release the residual head buffers. Cache initialization policy Given that our two head buffers persist across syscalls of the same thread, the next question is when to allocate such buffers. An option would be to simply allocate the head buffers at thread RESCALE – PU - Public – Page 23 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) (i.e., task struct) creation time. However, as shown in Figure 4, syscalls rarely fetch user data, e.g., ≈64% of the syscalls do not issue any fetches. Hence, eager head buffer allocation at thread creation time may lead to unnecessary memory overhead. As a result, the backend allocates the two head buffers (and the two regions) lazily, on the first user fetch that a thread incurs. Zero-copy optimization Given that most syscalls fetch little data, storing data into the cache is generally inexpensive. However, some syscalls are fetch-heavy and copy large amounts of data. In such scenarios, copying data into the cache incurs a nontrivial cost (as shown later in our evaluation), so it is desirable to eliminate as many such extra copies as possible. To investigate optimization avenues, we isolated all syscalls in our profiling results that copy more than a page of data from user space. Figure 8 presents our results in a heatmap. As shown in the figure, many of these syscalls (with the exception of exec) are I/O bound. To avoid resource contention, I/O bound syscalls rely on the kernel iov iterator functionality to copy dozens of large chunks of user data into kernel staging buffers. Such buffers (and their copied data) persist unchanged until the hardware (e.g., network card, hard disk) becomes available to complete the I/O transfer. To optimize out expensive copies to the cache, we use the kernel staging buffer itself as a data cache for all the iov iterator copies fetching one or more pages. To delay deallocation of such staging buffers until syscall exit, we increase the reference count on the underlying pages. When enabling this zero-copy optimization (SafeFetch default), our cache invalidation policy also releases the staging buffers on syscall exit. 3.7 Interconnection/Interfaces SafeFetch can be used as a standalone tool to harden a target Linux kernel. To build a SafeFetchenabled kernel, one can execute the following commands: $git clone git@github . com : vusec / safefetch . git safefetch $make -C safefetch mrproper $make -C safefetch x86_64_defconfig $cd safefetch && ./ scripts / config -d CONFIG_AUDITSYSCALL -e CONFIG_SAFEFETCH && cd .. $make -C safefetch $sudo make -C safefetch install RESCALE – PU - Public – Page 24 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) To enable SafeFetch at run time, one can execute the following commands: # enable syscall hooks for metadata /data management $./ safefetch_control .sh -hooks # enable safefetch in adaptive mode $./ safefetch_control .sh -adaptive 4096 4096 0 # disable safefetch $./safefetch_control.sh # enable safefetch in rbtree mode $./ safefetch_control .sh -rbtree 4096 4096 0 For more information, see https://github.com/vusec/safefetch. 3.8 Related Work While double-fetch or TOCTTOU (Time-of-Check-to-Time-of-Use) vulnerabilities affect different interfaces and components (e.g., enclaves [10,27], sandboxes [52], compilers [76], and compartments [5]), we focus here on closely related research on operating system kernels. Serna[64] was the first who coined the name “double-fetch vulnerability” to describe an instance in the Windows kernel. Since then, research has focused on finding double-fetch bugs (through static and dynamic program analysis) or even mitigating these issues. Static analysis On the static analysis front, Wang et al.[71] leverages pattern matching analysis on source code to find double-fetch bugs in the Linux kernel. DFTinker[49] extends such pattern-based approach to increase double-fetch coverage and reduce false positives. While pattern-matching techniques proved effective in uncovering new double-fetch bugs, they still produce a high rate of false positives and negatives and are fundamentally limited to detecting only specific bug patterns. DEADLINE[78] and DFTracker[72] improve static detection of double fetches by means of compiler-level symbolic execution. While these approaches can be applied more broadly (e.g., to detect compiler-induced double fetches[75,77]) and are generally better suited for vetting false positives, symbolic execution introduces other limitations (e.g., path explosion) and does not completely remove false reporting (e.g., due to imprecise memory modeling or incomplete code coverage). DEADLINE, for example, does not detect double fetches in inline assembly, which is widely used in the kernel. RESCALE – PU - Public – Page 25 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) HA HV Speculate KernelUser 3 4 HA HV call [rax] BTB Collision legit target eBPF gadget 1 2 call [rax] CA CV TA TV eBPF victim Figure 9: The BHI attack. The attacker first triggers the branch CAwith history HA1 , which inserts the target TAinto the BTB 2 , then triggers the victim branch CVwith history HV3 . The histories are crafted so that CAand CVshare the same BTB entry, so from CVthe CPU speculates to TA4 . 5.1.3 Branch History Injection In 2022, Branch History Injection (BHI) [3] showed that, despite mitigations, cross-privilege Spectre v2 is still possible on latest Intel CPUs by poisoning the Branch History Buffer (BHB). Figure 9 provides a high-level overview of the attack. In summary, by executing a sequence of conditional branches (HAand HV) right before performing a system call, an unprivileged attacker can cause the CPU to transiently jump to a chosen target (TA) when speculating over an indirect call in the kernel (CV). This happens because the CPU picks the speculative target for CVfrom a shared structure, the BTB, that is indexed using both the address of the instruction and the history of previous conditional branches, which is stored in the Branch History Buffer (BHB). Finding the right combination of histories that will result in a collision can be done with brute-forcing. To ensure the injected target, TA, contains a disclosure gadget, the original BHI attack relied on the presence of the extended Berkeley Packet Filter (eBPF), through which an unprivileged user can craft code that lives in the kernel. 5.1.4 Defenses As a recommended mitigation for BHI, Intel has advised to disable unprivileged eBPF, which is now disabled by default in the Linux kernel, and to mitigate potential disclosure gadgets by prepending a LFENCE instruction to them[18]. RESCALE – PU - Public – Page 32 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) Attack Surface Analysis The original BHI paper presented an initial estimate of the BHI attack surface beyond eBPF using simple data-flow analysis and a loose definition of disclosure gadget [3]. The analysis pinpointed 1,177 potential gadgets in the Linux kernel, but with no insights into their exploitability. Later, Intel researchers statically analyzed the Linux kernel [37], adopting a more refined data-flow-based approach and finding an order of magnitude more potential gadgets. Again, with the analysis unable to automatically reason over exploitability, Intel researchers resorted to manually assessing exploitability of the 8 simplest (linear) gadgets. No gadgets were deemed exploitable (6 due to reachability issues, 2 due to leakage constraints). Ultimately, with many potential gadgets uncovered by both scanners but no evidence of practical exploitability, no additional mitigations were deployed. Hardware mitigations On the hardware front, Intel has proposed a hardware mitigation, the BHI DIS S indirect predictor control, which prevents the CPU from selecting BTB entries based on history coming from lower security domains. As the performance overhead is nontrivial, future CPUs might come with an optimized version, namely BHI NO. At the time of writing, while Alder Lake and Raptor Lake Intel CPUs support BHI DIS S after applying a microcode update, the Linux Kernel does not have support to enable this feature. Advanced software mitigations Additional software mitigations, like Retpoline [17] or a software BHB-clearing sequence [18], have been proposed as spot fixes. However, they come with a prohibitive performance cost, discouraging practical deployment. For Linux, in particular, the software BHB-clearing sequence recommended by Intel has not been implemented at all, as developers rely on unprivileged eBPF being “the only known real-world BHB attack vector” [26]. IBT Indirect Branch Tracking (IBT) [66] is a defense for code-reuse attacks, such as return-oriented programming [65] and jump-oriented programming [9], which ensures that indirect branches always jump to an intended target. Intel CPUs implement a coarse-grained version of IBT in hardware, which ensures that every indirect branch lands on a special instruction (endbr32 or endbr64), which has to be inserted by the compiler. While IBT was originally designed to address architectural control-flow hijacking, it is also part of Intel’s mitigation guidance for BHI [18]. This is to provide defense-in-depth against speculative control-flow hijacking, limiting—somewhat similarly to the existing eIBRS—the possible Spectre-v2 disclosure gadget locations to the beginning of any indirect branch target in the kernel. Support for IBT was added to Intel processors with the Tiger Lake series [14] and is enabled by default from Linux kernel v6.2 [13]. RESCALE – PU - Public – Page 33 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) FineIBT Researchers have recently proposed a finer-grained IBT variant in software, called FineIBT [31]. The idea is to instrument the caller of each IBT-guarded indirect call to load a unique value into a register as well as the callee to check said value. If the value is different from the expected one, the execution path is directed to an illegal instruction to abort execution. This further restricts indirect calls to (architecturally or speculatively) target only compliant callees. Support for FineIBT was recently introduced in Linux kernel v6.2 [80]. Since FineIBT relies on Clang-instrumented kernels [31], to our knowledge, it is not yet enabled by default in any Linux distribution (unlike its IBT building block). 5.2 Threat Model We consider a traditional cross-privilege Spectre-v2 threat model, with a local unprivileged attacker seeking to disclose information from a privileged victim, such as the operating system kernel or the hypervisor. We specifically focus on a victim Linux kernel, running with all default Spectre-v2 mitigations on last-generation Intel CPUs such as eIBRS [16] and privileged eBPF. Finally, we assume other classes of vulnerabilities (e.g., memory errors) are subject of orthogonal mitigations and not part of the attack surface under study. 5.3 Overview To perform a Spectre-v2 attack against the kernel on last-generation Intel systems, one must target an indirect branch in the kernel that jumps to a disclosure gadget. Using BHI, one can then speculatively hijack a victim branch to the chosen gadget. InSpectre Gadget aids the analyst in choosing a suitable gadget, following the workflow depicted in Figure 10. As shown in the figure, the analyst provides InSpectre Gadget with a kernel image and a list of candidate gadgets in input. Our tool inspects each candidate for a fixed number of basic blocks with symbolic execution and returns a list of gadgets that lead to the transmission of a secret. Along with the gadgets, our tool outputs a number of gadget characteristics: the advanced exploitation techniques required (if any), the constraints that have to be met, the registers that have to be controlled, and the values that can be leaked. Such characteristics are stored into a database, which the analyst can later filter according to what targets are reachable, what registers are controlled when the speculative hijack happens, and what mitigations are enabled for a given target. The resulting gadgets can then be exploited to mount end-to-end Spectre-v2 kernel attacks. In the next sections, we first elaborate on how InSpectre Gadget models gadgets (constraints, exploitability, etc.) and how it then performs exploitation-aware gadget inspection. RESCALE – PU - Public – Page 34 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) Exploitable Gadgets Filtered Gadgets 1 2 3 Control Defenses Craft Exploit call [rax] call [rax] InSpectre Gadget Chosen gadget + Kernel Targets Figure 10: InSpectre gadget workflow. The analyst provides a kernel image and a list of target addresses to InSpectre Gadget 1 , which performs in-depth inspection to find gadgets that can leak secrets and output their characteristics. The gadgets can be filtered 2 based on the available attacker-controlled registers and the mitigations enabled, and used to craft Spectre-v2 exploits against the kernel 3 . 5.4 InSpectre Gadget A recurring problem in transient execution attacks is evaluating if a given instruction sequence can leak a secret via a—typically cache [79,47,34,1,55,42]—covert channel. To leak a secret through the cache, an attacker needs to open a speculation window and accommodate adisclosure gadget. In this section, we explain that practical disclosure gadgets extend well beyond those that fit existing narrow definitions. Next, we show how InSpectre Gadget uses symbolic execution to analyze the candidate gadgets for Spectre-v2 exploitability. 5.4.1 Standard Gadgets Today’s tools typically concentrate on finding speculation windows and overapproximating exploitability—leaving the analysis of disclosure gadgets to the human analyst. Unfortunately, the complexity of such exploitability analysis is daunting and, unsurprisingly, analysts generally look for standard Spectre disclosure gadgets, such as the one in Listing 1. 1// Load from attacker - controlled address 2uint64_t secret = * attacker ; 3// Mask the loaded value 4uint8_t secretByte = ( secret | & 0 xFF); 5// Shift the result RESCALE – PU - Public – Page 35 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) 6uint32_t tsecret = secretByte << 9; 7// Use transmission secret as index . 8uint64_t transmission = *( tbase + tsecret ); Listing 1: Standard Spectre disclosure gadget. In a standard (or “perfect”) gadget, the CPU loads a secret from an attacker-controlled address. The loaded secret must be sufficiently small, either by nature, or as the result of a bitmask (Line 4), to serve as an index into a second buffer—ideally shared with the attacker. This second buffer is known as the reload buffer. Moreover, to ensure that each value of the secret corresponds to a different cache line, the gadget should shift the value left by some stride (Line 6). The result is known as the transmitted secret. As the gadget subsequently adds it to a second attacker-controlled value which we refer to as the transmission base and dereferences the resulting value, it inadvertently establishes a transmission through a cache covert channel. Specifically, by iterating over the reload buffer with the same stride while timing the accesses, attackers can infer that the secret value is the index of the buffer element for which the access is fast (because it is in the cache). 5.4.2 Exploitation-Aware Gadget Analysis Standard gadgets are the most intuitive to understand and the most straightforward to exploit, but they are by no means the only exploitable ones. While limiting the analysis to standard gadgets reduces the complexity, attackers are under no obligation to respect such restrictions. BLINDSIDE [32], KASPER [38], PACMAN [59], and RETBLEED [74], for example, all exploit gadgets that deviate from such narrow definitions. Moreover, besides leaking information through the cache, attackers may avail themselves of a myriad of other covert channels [33,62,29,7,61]. In this section, we relax the assumptions for standard gadgets, describe the additional challenge posed by each relaxation, and where appropriate, explain how we can meet the challenge and still leak secrets. C1. Base not controlled To perform FLUSH+RELOAD an attacker needs to be in control not only of the secret address, but also of the transmission base, so that the transmission load falls inside of the reload buffer. However, if the base is not controlled, attackers can still perform PRIME+PROBE [47]. C2. Secret entropy too big If the code does not mask the secret prior to its transmission, the entropy impedes its recovery by the attacker who would have to probe too many memory locations. However, if the attackers control the transmission base, they can still exploit such gadgets by repurposing a technique pioneered in earlier attacks [32,62,74], which we call the known-prefix technique. First, the attacker chooses a location in memory near the secret that contains a small known value, and uses that as the secret address. Then, by increasingly shifting the address, the attacker makes RESCALE – PU - Public – Page 36 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) buff - 0xdeadbe00 + 0xdeadbeef Transmission Base Known Prefix 00 00 00 00 de ? 00 00 00 00 de ad be ? 00 00 00 00 ? 1 2 3 Secret Figure 11: Known-prefix technique. The attacker first points the secret address to some known data 1 . Then, by shifting the address 2 , small portions of the secret are revealed. With that, one can adjust the base of subsequent transmissions 3 . sure that only a few bytes of the secret are unknown at any given transmission, and adjusts the transmission base according to the bytes that are already known (see Figure 11). C3. Max secret too high Another problem that may occur when leaking a large secret value is that the transmission address might end up outside the valid address space. However, if enough bits of the transmission base are under attacker control, one can adjust the base to overflow the secret value, ensuring that the transmission always occurs in the valid address space. We call this technique base adjusting and it is often used in conjunction with the known-prefix technique. C4. Secret too small If the gadget does not shift the secret value prior to transmission, the lower bits cannot be recovered through cache covert channels, since nearby addresses belong to the same cache line. However, if at least one byte of the secret is above cache-line granularity, we can use the knownprefix technique to leak the uppermost byte in each iteration. Otherwise, an attacker may use the sliding technique of RETBLEED [74], which exploits the fact that prefetchers generally do not prefetch cache lines across pages [67]. An attacker may now adjust the base address to be near a page boundary, so that even a one-bit difference in the secret value will result in the transmission being observed on another memory page. Doing so eliminates any cache and prefetcher noise that would otherwise make two different values indistinguishable. RESCALE – PU - Public – Page 37 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) C5. Base aliasing In some cases, even if the transmission base is attacker-controlled, the base cannot be chosen without influencing the secret address—a phenomenon we refer to as aliasing. This may occur, for instance, if the base is computed from a value that is also used to compute the secret address. In this case, an attacker can still leak values using PRIME+PROBE. However, not all secret addresses may be targeted as the probe region must reside within mapped memory. Specifically, since the probe region typically follows the secret address, the final secret addresses result in a non-mapped probe region. Other, more complex dependencies between the base and the secret address out of scope. InSpectre Gadget marks these cases as not (easily) exploitable. C6. Non-linear gadgets An implicit assumption of most static analysis techniques for gadget scanning is that all the instructions have to be in the same basic block. Since modern CPUs all support nested speculation, this is not a limitation in practice. The attacker can either train branches or simply use static prediction to ensure that the transmission gadget is reached during speculative execution. An attacker can also use SMT contention, as demonstrated by prior work [51], to delay branch resolution and create large speculation windows that can accommodate multiple basic blocks. C7. Other transmitters Finally, a plethora of covert channels exist in modern CPUs outside of caches (e.g., the TLB [48]) and one should flexibly support different transmitters. For brevity, we focus our main analysis on classic secret-dependent data load/store transmitters and later show our tool can be easily extended to support other vectors such as the recent SLAM covert channel [36]. 5.4.3 Design While advanced exploitation techniques relax the assumptions for standard gadgets, each has its own requirements and constraints for applicability. To reason over the constraints and to assess if the code satisfies them, InSpectre Gadget employs symbolic execution—expressing variables as symbols, and exploring all the possible directions of a program’s control flow at the same time, while recording the corresponding symbolic constraints. To this end, we build our tool on top of ANGR[68], a symbolic execution engine widely used in the field of software security and reverse engineering. Although symbolic execution quickly leads to state explosion in the general case, InSpectre Gadget explores only a small fraction of the program’s control flow. In particular, since the number of instructions executed during a speculation window is limited, symbolic execution can easily explore multiple paths, while keeping the number of explored states small. All the steps described below are performed automatically by InSpectre Gadget. RESCALE – PU - Public – Page 38 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) Tracking attacker control We start our analysis by substituting all the values stored in registers and on the stack with symbolic variables, which for now we assume are attacker-controlled, and marked as such. Next, we symbolically execute code starting from a given location. On each load, we produce a new symbolic value with a label attached to it. If the load comes from an attacker-controlled symbol, it is marked as a potential secret. If the address comes from a potential secret, it is marked as a potential transmission. Note that potential secret implies attacker-controlled, since it is any value loaded from an attacker-chosen location. Similarly, potential transmission implies potential secret. This strategy allows us to track of attacker control through complex chains of loads. Store-to-Load Forwarding We model Store-To-Load Forwarding by keeping a list of all the symbolic stores, and checking this list for each of the loads encountered by the symbolic execution engine. If the symbolic expression of the load address aliases with that of a previous store, we forward the stored value to the load. Store-to-Load forwarding can be enabled or disabled with a runtime flag. Potential transmissions We let the symbolic execution engine run for a configurable number of basic blocks, recording all the constraints that might be added, for instance by cmove or branch instructions. We also record the symbolic expression of each load. Finally, after the scanning phase is finished, the scanner reports a list of potential transmissions, i.e., loads whose symbolic address has been marked as a potential secret. In a second phase, we inspect the AST of the symbolic expression of each potential transmission, identifying the transmission base, transmitted secret, and secret address. A high-level example is shown in Figure 12. Finally, we check if the base depends on any value used to construct the secret address, perform a range analysis to infer the minimum, maximum and stride of all the transmission components, and perform an inferable-bits analysis to infer which bits of the secret end up in the transmission and at which position. This comprehensive analysis is crucial for accurate exploitability reasoning. For instance, data-flow information alone is insufficient to determine the controllability requirements to be met. Gadget reasoning With this information, we can finally reason about each gadget. All the properties found during the analysis, along with a list of registers that the attacker needs to control for each gadget, are saved in a database. We now use a reasoner to model exploitation techniques with database queries. The reasoner inserts columns indicating which gadgets can leak a secret, and, if needed, which techniques are required. For instance, if a gadget has a high secret entropy (i.e., number of transmitted bits >16), the reasoner checks if we can perform the known-prefix technique (i.e., secret address is attacker-controlled with granularity <=16 bits). RESCALE – PU - Public – Page 39 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) Transmission Base Secret Transmitted Secret LOAD[ LOAD[rax] + (( LOAD[rbx] & 0xff) * 8) ] Figure 12: Anatomy of a transmission. Once a potential transmission is identified through symbolic execution, InSpectre Gadget dissects its symbolic expression into components, which are then analyzed to reason about exploitability. 5.4.4 Limitations As opposed to tools like KASPER [38], InSpectre Gadget is designed to analyze only the content of speculation windows (e.g., call and jump targets for Spectre-v2), whose entry points have to be provided by the analyst. Other aspects needed for end-to-end exploitation, such as the reachability of the target and the presence of a suitable victim branch, are not part of the tool’s output. Moreover, InSpectre Gadget cannot completely prove the absence of gadgets in a given snippet of code. Regarding exploitability results, our current prototype has a number of limitations potentially impacting accuracy. First, our tool is based on ANGR and relies on both its disassembler (CAPSTONE) and constraint solver (Z3). Whenever an error occurs in one of these components, e.g., on unsupported instructions, we have to bail out from the analysis, leaving some symbolic states unexplored. Second, transmissions that contain a complex symbolic expression, e.g., two independently-controlled loads used in a XOR operation, cannot be easily unpacked into a base and a transmitted secret, but might still leak a value. We mark these cases as complex and approximate their ranges by querying the SAT solver for the minimum and maximum values of the whole expression and of each sub-expression. We also use this approach when performing range analysis on expressions with complex (symbolic) constraints, whose values cannot be easily reduced to an interval or a small set. For these cases, besides reporting the minimum and the maximum values, we also check if certain bits are always 0 or 1 to approximate the stride. Finally, at the moment our tool models the attacker’s control over complex chains of loads as a binary condition (controlled or not controlled), as opposed to reporting the degree of control as done for transmission components. For complex aliasing cases, this can introduce imprecision (e.g., inaccurate classification of the required exploitation techniques). 5.5 Interconnection/Interfaces InSpectre Gadget can be used as a standalone tool—or easily integrate with other tools and workflows—in order to analyze the Spectre attack surface of an input program binary. A typical workflow operates as follows: RESCALE – PU - Public – Page 40 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) # Analyze <CSV > targets for binary <BINARY > $inspectre analyze <BINARY > --address -list <CSV > -- config config_all .yaml --output out gadgets . csv --asm out/asm # Perform exploitability analysis of any gadgets found $inspectre reason out / gadgets . csv out /gadgets - reasoned . csv The workflow above generates a report of the form: [-] Imported 32 gadgets [-] Performing exploitability analysis ... Found 2 exploitable gadgets ! [-] Saving to out /gadgets - reasoned . csv [-] Done! For more information, see https://github.com/vusec/inspectre-gadget. 5.6 Related Work Spectre gadget scanners Spectre gadget scanners documented in literature mostly focus on Spectre v1. With exceptions [58], such scanners typically rely on dynamic analysis. SPECFUZZ [54] uses fuzzing to detect out-of-bounds accesses on a speculative path. SpecFuzz marks every out-of-bound access as a gadget, without modeling attacker controllability. In response, SPECTAINT [56] requires the secret address to be tainted with attacker input, as determined via dynamic taint analysis (DTA). KASPER [38] also relies on DTA for gadget characterization, but generalizes the fixed patterns used by SpecTaint. Unlike InSpectre Gadget, all these solutions identify gadgets solely based on their data flow, an overapproximation that leaves their exploitability uncertain. Like ours, other gadget scanners are based on symbolic execution, but typically focus on other use cases (i.e., verification [35] or early detection [70]) with no exploitability analysis. Spectre-v2 attacks Besides work exclusively focusing on gadget scanning, prior Spectre-v2 attack efforts also described gadget analysis campaigns. For instance, the RETBLEED authors [74] used static data-flow analysis to identify basic 3-load gadgets for a native Spectre-RSB-to-BTI exploit. This simple strategy is sufficient as Retbleed exploits (similar to Inception [69]) vulnerable older-generation CPUs with no eIBRS. As such, they can speculatively hijack control flow to arbitrary code locations with no restriction. The BHI authors [3], also using data-flow-based gadget analysis, presented evidence eBPF=off exploits were at least potentially feasible on modern eIBRS-enabled platforms—but with no RESCALE – PU - Public – Page 41 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) [47] Fangfei Liu, Yuval Yarom, Qian Ge, Gernot Heiser, and Ruby B. Lee. Last-level cache side-channel attacks are practical. In IEEE S&P, 2015. [48] Kevin Loughlin, Ian Neal, Jiacheng Ma, Elisa Tsai, Ofir Weisse, Satish Narayanasamy, and Baris Kasikci. DOLMA: Securing speculation with the principle of transient nonobservability. In USENIX Security, 2021. [49] Yingqi Luo, Pengfei Wang, Xu Zhou, and Kai Lu. Dftinker: Detecting and fixing doublefetch bugs in an automated way. In WASA, 2018. [50] Giorgi Maisuradze and Christian Rossow. ret2spec: Speculative execution using return stack buffers. In CCS, 2018. [51] Alyssa Milburn, Ke Sun, and Henrique Kawakami. You cannot always win the race: Analyzing mitigations for branch target prediction attacks. In IEEE EuroS&P, 2023. [52] Shravan Narayan, Craig Disselkoen, Tal Garfinkel, Nathan Froyd, Eric Rahm, Sorin Lerner, Hovav Shacham, and Deian Stefan. Retrofitting fine grain isolation in the Firefox renderer. In USENIX Security, 2020. [53] National Security Agency. Ghidra. [54] Oleksii Oleksenko, Bohdan Trach, Mark Silberstein, and Christof Fetzer. SpecFuzz: Bringing Spectre-type vulnerabilities to the surface. In USENIX Security, 2020. [55] Colin Percival. CACHE MISSING FOR FUN AND PROFIT. Proceedings of BSDCan 2005, 2005. [56] Zhenxiao Qi, Qian Feng, Yueqiang Cheng, Mengjia Yan, Peng Li, Heng Yin, and Tao Wei. SpecTaint: Speculative taint analysis for discovering Spectre gadgets. In NDSS, 2021. [57] Radare org. Radare2. [58] Hany Ragab, Andrea Mambretti, Anil Kurmus, and Cristiano Giuffrida. GhostRace: Exploiting and mitigating speculative race conditions. In USENIX Security, 2024. [59] Joseph Ravichandran, Weon Taek Na, Jay Lang, and Mengjia Yan. PACMAN: Attacking ARM pointer authentication with speculative execution. In ISCA, 2022. [60] ReFirm Labs. Binwalk. [61] Xida Ren, Logan Moody, Mohammadkazem Taram, Matthew Jordan, Dean M. Tullsen, and Ashish Venkat. I see dead µops: Leaking secrets via Intel/AMD micro-op caches. In ISCA, 2021. [62] Stephan Van Schaik, Alyssa Milburn, Sebastian ¨ Osterlund, Pietro Frigo, Giorgi Maisuradze, Kaveh Razavi, Herbert Bos, and Cristiano Giuffrida. RIDL: Rogue in-flight data load. In IEEE S&P, 2019. [63] Michael Schwarz, Daniel Gruss, Moritz Lipp, Cl´ ementine Maurice, Thomas Schuster, Anders Fogh, and Stefan Mangard. Automated detection, exploitation, and elimination of double-fetch bugs using modern cpu features. In AsiaCCS, 2018. RESCALE – PU - Public – Page 48 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) [64] Fermin J. Serna. MS08-061 : The case of the kernel mode double-fetch. https://msrc.microsoft.com/blog/2008/10/ ms08-061-the-case-of-the-kernel-mode-double-fetch/, 2008. [65] Hovav Shacham. The geometry of innocent flesh on the bone: Return-into-libc without function calls (on the x86). In CCS, 2007. [66] Vedvyas Shanbhogue, Deepak Gupta, and Ravi Sahita. Security analysis of processor instruction set architecture for enforcing control-flow integrity. In S&P, 2019. [67] Youngjoo Shin, Hyung Chan Kim, Dokeun Kwon, Ji Hoon Jeong, and Junbeom Hur. Unveiling hardware-based data prefetcher, a hidden source of information leakage. In ACM CCS, 2018. [68] Yan Shoshitaishvili, Ruoyu Wang, Christopher Salls, Nick Stephens, Mario Polino, Andrew Dutcher, John Grosen, Siji Feng, Christophe Hauser, Christopher Kruegel, and Giovanni Vigna. (state of) the art of war: Offensive techniques in binary analysis. In IEEE S&P, 2016. [69] Dani¨ el Trujillo, Johannes Wikner, and Kaveh Razavi. Inception: Exposing new attack surfaces with training in transient execution. In USENIX Security, 2023. [70] Guanhua Wang, Sudipta Chattopadhyay, Arnab Kumar Biswas, Tulika Mitra, and Abhik Roychoudhury. KLEESPECTRE: Detecting information leakage through speculative cache attacks via symbolic execution. ACM TOSEM, 2020. [71] Pengfei Wang, Jens Krinke, Kai Lu, Gen Li, and Steve Dodier-Lazaro. How double-fetch situations turn into double-fetch vulnerabilities: A study of double fetches in the Linux kernel. In USENIX Security, 2017. [72] Pengfei Wang, Kai Lu, Gen Li, and Xu Zhou. Dftracker: Detecting double-fetch bugs by multi-taint parallel tracking. Frontiers of Computer Science, 2019. [73] Sander Wiebing, Alvise de Faveri Tron, Herbert Bos, and Cristiano Giuffrida. InSpectre Gadget: Inspecting the Residual Attack Surface of Cross-privilege Spectre v2. In USENIX Security, 2024. [74] Johannes Wikner and Kaveh Razavi. RETBLEED: Arbitrary speculative code execution with return instructions. In USENIX Security, 2022. [75] Felix Wilhelm. Xenpwn: Breaking paravirtualized devices. Black Hat USA, 2016. [76] Jianhao Xu, Luca Di Bartolomeo, Flavio Toffalini, Bing Mao, and Mathias Payer. WarpAttack: Bypassing CFI through compiler-introduced double-fetches. In IEEE S&P, 2023. [77] Jianhao Xu, Kangjie Lu, Zhengjie Du, Zhu Ding, Linke Li, Qiushi Wu, Mathias Payer, and Bing Mao. Silent bugs matter: A study of compiler-introduced security bugs. In USENIX Security, 2023. [78] Meng Xu, Chenxiong Qian, Kangjie Lu, Michael Backes, and Taesoo Kim. Precise and scalable detection of double-fetch bugs in OS kernels. In IEEE S&P, 2018. RESCALE – PU - Public – Page 49 / 50
D3.7: Solutions for Security Testing of Low-Level System Components (first version) [79] Yuval Yarom and Katrina Falkner. FLUSH+RELOAD: A high resolution, low noise, L3 cache side-channel attack. In USENIX Security, 2014. [80] Peter Zijlstra and Joao Moreira. Implement fineibt. https://lore.kernel.org/lkml/ 20221018133550.1160-1-[email protected]/, 2022. RESCALE – PU - Public – Page 50 / 50