Pint, herramienta de simulación basada en trazas Pin
Abstract
In the course of this project we have developed a set of programs to improve the correction and execution time of the gem5 simulator. For this, we moved the functional simulation step out of gem5 into an independent instrumented process to ensure correction in the functional stage and to provide a good execution speed (since the code will then be natively executed). This instrumentation is done by Pin. Also, in order to allow efficient communication between the processes despite the limitations imposed by Pin to the available tools, an IPC framework to allow message passing between the processes was developed. This framework uses lockless fifo queues over shared memory so the resulting slowdown is minimal.
Full text
Escola Tècnica Superior d’Enginyeria Informàtica Universitat Politècnica de València PINT, HERRAMIENTA DE SIMULACIÓN BASADA EN TRAZAS PIN Final Year Project Computer engineering Author: Francisco Blas Izquierdo Riera Director: Julio Sahuquillo Borrás - UPV Per Stenström - CTH
December 18, 2012 2
Abstract In the course of this project we have developed a set of programs to improve the correction and execution time of the gem5 simulator. For this, we moved the functional simulation step out of gem5 into an independent instrumented process to ensure correction in the functional stage and to provide a good execution speed (since the code will then be natively executed). This instrumentation is done by Pin. Also, in order to allow efficient communication between the processes despite the limitations imposed by Pin to the available tools, an IPC framework to allow message passing between the processes was developed. This framework uses lockless fifo queues over shared memory so the resulting slowdown is minimal. Keywords: hardware, simulator, x86, Pin, Pintool, gem5, ipc, fifo, C++
Contents 1 Introduction 4 1.1 Projectrationale ......................... 4 1.2 Project objectives . . . . . . . . . . . . . . . . . . . . . . . . . 5 1.3 Projectstrengths ......................... 5 1.4 Memorystructure......................... 7 2 State of the art 8 2.1 Architectural simulators . . . . . . . . . . . . . . . . . . . . . 8 2.1.1 Graphite.......................... 10 2.1.2 Multi2Sim......................... 10 2.1.3 gem5............................ 11 2.2 Instrumentation systems . . . . . . . . . . . . . . . . . . . . . 11 2.2.1 gprof............................ 12 2.2.2 Pin............................. 12 2.3 Benchmarks............................ 13 2.3.1 SPECCPU2006...................... 13 2.3.2 SPLASH-2......................... 14 3 System description 16 3.1 Mead: a message passing framework . . . . . . . . . . . . . . 16 2
Contents Contents 3.2 Pint: a Pin based trace generator . . . . . . . . . . . . . . . . 18 3.3 Schnapps: a simple consumer of the traces . . . . . . . . . . . 19 3.4 Gin5: a gem5 trace player . . . . . . . . . . . . . . . . . . . . 20 4 System design 21 4.1 Mead: a message passing framework . . . . . . . . . . . . . . 21 4.2 Pint: a Pin based trace generator . . . . . . . . . . . . . . . . 23 4.3 Schnapps: a simple consumer of the traces . . . . . . . . . . . 25 4.4 Gin5: a gem5 trace player . . . . . . . . . . . . . . . . . . . . 26 5 Results 28 6 Conclusions 30 6.1 Improvements for next release . . . . . . . . . . . . . . . . . . 31 A User manual 32 A.1 Building.............................. 32 A.2 Pint ................................ 32 A.3 Schnapps ............................. 33 A.4 Gin5................................ 33 B Relevant source code 34 Bibliography 67 3
Chapter 1 Introduction 1.1 Project rationale Despite the vast amount of hardware simulators that exist nowadays, most of them either lack flexibility on the simulations or are slow since they simulate the code execution instead of instrumenting the natively executed executable. As a related issue, since code is simulated and not executed it is common to find bugs where the simulator will not set the processor state properly which cause corner cases where not acting as the processor causes execution issues with some programs. Also many simulators lack support for parallel execution and those who do tend to add big overheads when running the simulation in a single machine and will not support instruction level simulation granularity. Finally current simulators tend to add big overheads to the functional simulation step which makes it unfeasible to run large tests even when simulating simple systems. Program instrumentation solves all these shortcomings by running the code natively (modified so it will also execute the instrumentation code), 4
Chapter 1. Introduction 1.2. Project objectives allowing it to run in parallel and, since code is executed natively in the processor, providing a completely native execution. Given the limitations of the current simulators we consider that the community needs a flexible and fast instrumentation based tracer able to be used with a broad range of programming languages who will handle local simulations in parallel, with small overheads and with instruction level granularity. 1.2 Project objectives Our main objective is providing an instrumentation based tracer that can be used with other simulators. Given the problems with the size these traces can have we will feed them in a lively fashion. In order to see whether these objectives are met or not we will measure the slowdown compared to the non instrumented program with a simple trace consumer (to ensure it is not the bottleneck). Our objective is getting at least similar slowdowns to the ones of Graphite [9], but removing the caveats it has at least on single core processors. Also we intend to design an architecture which can later be expanded to support multiple simultaneous execution threads. This will be done on a later version of the project though due to timing constraints. 1.3 Project strengths The biggest problem with instrumentation-based systems is that the instrumentation code is limited heavily by the instrumentation API of the instrumentation system (for example the POSIX thread API can not be used with Pin), this also reduces vastly the number of languages that can be used, in 5
1.3. Project strengths Chapter 1. Introduction order to overcome these limitations, we use FIFO queues placed in shared memory to extract the data from a process to another, using the operating system process separation to execute the simulation in a different processor. As a result, the simulator overcomes the restrictions caused by the instrumentation framework since these will only apply to the process where the program is being instrumented. As a side effect of this approach, the resulting simulators will be segmented since the functional simulation can be done on a different processing unit than the one running the simulation itself. As the number of processor cores increases it is likely that hardware simulators will use the segmentation approach more extensively in order to increase performance. Also, when shared the cache accesses caused by the shared memory communication cause slowdowns, hyperthreading processors can be used and proper processor affinity to the processes can be set so the critical simulation parts (i.e those responsible of bottlenecks) will be set along with the previous part on the different threads of a single core so the reads from the FIFO queue are likely to be on the level 1 cache. To ensure real parallelism each thread of the instrumented program can use a different FIFO queue to extract its traces1. This also allows the user to limit instruction granularity by setting an appropriate queue size since the instrumentation will stop execution once the queue is full2. As an example of how Pint can be used we also created Gin5 a slightly modified version of the gem5 simulator which uses Pint’s instrumentation as the source of the memory access information during the simulation. Another known problem is that simulators tend to be very good on sim1The completely multithreaded implementation will be finished with the next release of Pint 2For this a simulation started event is added to ensure the first instruction will not be executed until the simulator wants it. 6
Chapter 1. Introduction 1.4. Memory structure ulating a specific part of the system whilst having issues on others. We consider that in the future this FIFO system may be useful to interconnect simulators so the best of them can be gotten. 1.4 Memory structure In this introduction we presented the problem we are trying to fix, our objectives and our strengths. On the next section, we will explain the state of the art at the time of our publication in the topics of Architectural Simulators, Instrumentation systems and Benchmarks. Afterwards we will analyze the four modules we developed for the project and we will continue later with the design decisions. We will finally present our benchmarking results and our conclusions. Annexed you will find a brief user manual in case you want to try our system and the referred bibliography. On the Annex folder you will find the sources we developed in this project. 7
2.3. Benchmarks Chapter 2. State of the art system and with a main focus on the CPU execution speed. Despite being there since the 2006 these benchmarks are widely used and understood in the academic and real world. Most of the programs provided with the benchmark are licensed with GPL style licenses and are well known in the free software world, for example gcc or perl, whilst others come from different research projects. It is because of this that the copyright is held over the input files in this benchmarks. The main problem with these benchmarks is that they focus on single threaded processes. 2.3.2 SPLASH-2 The SPLASH-2 [23] [26] benchmark was developed by the Flash research group at the Stanford university to provide a set of benchmarks that could be used on shared memory multiprocessor systems. Although the benchmarks are quite old and require modifications to work properly they can still be used and have the advantadge of running in a short time. The applications provided are related to the scientific world with examples of 3 body gravity simulators or some kernels like the LU decomposition of a matrix. Since the original tests will not run, we used a modified version of the SPLASH-2 benchmark [19]. Even more modifications were required for the null macro to work properly and for the tests to be able to be run with Pin on hardened systems, these modifications are provided as a patch file in the source distribution. The main reason for choosing these was that the relative performance results of these tests (although with more than one processor) were provided on [9] so we did not need to run the benchmarks again for Graphite and thus 14
Chapter 2. State of the art 2.3. Benchmarks set up the required Debian environment. 15
Chapter 3 System description Our application will be divided in 4 modules: Mead, a framework for providing an efficient message passing interface between different processes; Pint, a Pin based trace generator; Schnapps, a simple consumer of the traces; and Gin5, a gem5 trace player for the memory system. The traces will be generated by Pint and then fed through Mead to either Schnapps or Gin5 which will process it and provide some simulation statistics. 3.1 Mead: a message passing framework The pattern of message passing is not new and it conforms the base of some Object Oriented views. Mead will provide a fast and simple way of passing around the traces as messages stating that something has happened (for example the program made an execution memory access of size x at position y). These messages may contain the thread identifier of the thread that caused them and also attached data, for example in the case of a memory write the data available before writing and the data being written. Although the API provided by mead is quite agnostic of the message 16
Chapter 3. System description 3.1. Mead: a message passing framework passing system being used we have chosen producer-consumer FIFO queues. FIFOs are used since they are a known pattern which allows for easy implementation and migration over other interprocedural communication systems, if interprocess shared memory is not an option, like POSIX message queues or datagram sockets. Our FIFO model differs slightly from the standard model since it allows for two communication types, on one hand you have the event communication system which can queue many events for further handling by the receiving side. On the other you will find a command interface able of holding a single command. The command interface requires acknowledging the sent command and is used to indicate important events which require specific handling by the queue system like the death of the FIFO or the beginning and ending of the simulation procedures. The main difference between events and commands are that events are unidirectional (from the producer to the consumer) whilst commands can be used bidirectionally (as long as the absence of collisions is guaranteed by the programmer) and are more easily handled with the futex syscall which makes them very useful for events which will require a really heavy processing on the other side by allowing the other thread to preempt the CPU while this is done. The FIFO architecture is based over a central FIFO (called the main FIFO) which is used to send global events which are supposed to stall the simulation until attended (so the simulator can decide whether it should clear or not the per thread queues before processing the aforementioned event), this queue handles at least the thread creation and deletion events where a new FIFO queue is negotiated between both sides, but it can also be used to process events like the creation and deletion of new mappings amongst 17
3.2. Pint: a Pin based trace generator Chapter 3. System description others. Given the importance of the main FIFO in the architecture it is important that both processes know where to access it beforehand and are able to negotiate its creation independently of who arrived first (since synchronization is impossible before the FIFO creation). The framework also features a per thread FIFO which can be used to send events which are not of global significance to the listener on the other side. This lets the programmer communicate information fast since the queues can be then lockless and, as a result, as long as there are at least two processors available the current process will not be changed by the kernel preventing expensive context switches. The creation of these FIFOs should ne negotiated over the main FIFO when implementing the multithreaded version. 3.2 Pint: a Pin based trace generator Pint by itself it is not a simulator but a framework providing efficient ways to extract the data from the instrumented program through Mead. The version presented with this project is single threaded (although designed to be multithreaded and with part of the work for that already done) and relies on Mead for communicating with the simulator itself. As an example the provided instrumentation will study memory accesses made by the program (of any type ranging from prefetches to execution fetches) and sends them out with Mead so the simulator can prevent the issues associated with Pin tracing tools. The code here is focused heavily on speed and thus the user must have the option to choose the features that should to be used. The granularity of the execution can be easily tunned by setting an ap18
Chapter 3. System description3.3. Schnapps: a simple consumer of the traces propriate queue size. For example for instruction by instruction execution the queue must have size one. Also, a simulation started event must be the first one to be queued so you can discard old elements when you want the execution to be done. Since all instructions will start (and contain) with a single fetch event it is possible to use this event as the differentiator between instructions. Anyway it is a good idea to integrate at least the number of events the instruction will cause to make tracing easier. This may be done on future versions. Pint also provides a way to specify the number of instructions that must be executed before switching to the next simulation mode and thus you can provide the number of instructions that must be executed (by the sum of threads) before switching to another simulation mode. The mode automaton allows for three simulation modes which are switched in the following order, the fast forward, the warm up and the simulation mode, which will then go back to the fast forward. In the fast forward mode instructions are just accounted and executed but no data is generated which allows for near native speed execution. In the warm up mode and simulation mode instructions will generate events for filling the caches but the entrance and exit of the simulation status are notified to the consumer so it can handle statistics properly. 3.3 Schnapps: a simple consumer of the traces Schnapps is intended to be used mainly for analyzing the performance of the instrumentation code by consuming the events generated whilst trying to avoid causing any bottlenecks in execution, and also as an example program of how to extract the generated traces. 19
3.4. Gin5: a gem5 trace player Chapter 3. System description Schnapps reads the traces generated by Pint and outputs the map changes as they happen (in a diff like format) and some statistics for the current simulation and for the total run, in particular, amount of data read or written by the different memory access types, the number of said accesses that has happened and an execution mark made by xoring the different accesses’ addresses together the idea being that different marks imply different traces being generated but the same mark does not necessarily imply the same trace being generated. . 3.4 Gin5: a gem5 trace player Although previous versions of gem5 came with a trace player supporting different formats, these modules stopped being maintained long time ago and those stopped building, as a result and despite being a good base for starting the work given the big amount of changes the memory system has suffered since then a different base was necessary. As a result we have set again a generic CPU for playing traces (missing the TLB) which will ask for memory access request to the queues via a clear interface so it can be used also with other types of trace formats including the old ones if the classes containing them are updated. 20
Chapter 4 System design 4.1 Mead: a message passing framework Mead has a macro of particular interest: USE_YIELD which will enable the use of the yield system call to let other process use the processor when waiting. On mead, we have chosen to implement a lockless buffer ring over shared memory for our FIFOs since it is a well known pattern [4] [24]. The other reasons for choosing such an structure was speed and independence. By being lockless we avoid expensive spinlocks that would hinder performance whilst avoiding also having to either use the ones provided by Pin everywhere or building our own. Also having the data in shared memory will prevent us from making expensive system calls to have the messages passed and will allow the usage of cache for that. It should be taken into account that the lockless queue will only work properly if a single thread acts as reader and a single thread acts as writer. In case of having more threads at either side they require a lock to work properly. 21
4.1. Mead: a message passing framework Chapter 4. System design Our implementation is based on templates so it can be used with different classes although it should be taken into mind that the same class (i.e. no inheritance) should be used on the whole queue. Also one of the current major caveats is that the shared memory address is currently hardcoded and, as a result, only a single instance of the program can be started at the same time. We expect to fix this in future versions by providing a launcher that will allocate an anonymous shared memory segment and pass its identifier to both Pint and the trace reader being used. The command types are defined by the shmstatus enum. Since newly allocated shared memory is filled with 0s we assume the 0 value as the initial state (NONE). The server will then write a SERVER_STARTED command and wait for a CLIENT_ACK then. Finally when dying the server is expected to send the SERVER_DIED command so the client will not wait forever for data. Of the many methods provided, those of special relevance for the programmer are the gethead and gettail methods used to be able to access the data we want to insert or extract from the queue, the push and pop methods used for adding or removing an element from the queue and the full and empty methods used to check for these states. Also some methods for waiting in case the queue is empty/full are provided but these must be used carefully since if the consumer is singlethreaded it could hang waiting forever for the queue to match the condition. In this case special waits monitoring the main queue status too are recommened instead. Also a wait_push method is provided that will wait until a push can be done. For control handling we provide the send_control, receive_control and ack_control methods. In particular send_control will wait until an ACK is 22
Chapter 4. System design 4.2. Pint: a Pin based trace generator sent back to state the condition was taken care of. Finally the wait_start and tell_Start methods are provided for initialization and instead of yields they use calls to the futex syscall to lock the thread until they have been attended to reduce the processor load in some situations. 4.2 Pint: a Pin based trace generator In order to ensure unwanted features will not hinder performance preprocessor based switches can be used to disable those you are not interested in using, also some other options can be set in this way. The macros of interest here are PADSIZE which defines the amount of bytes of the cache line in order to prevent false sharing, MAXMEMSIZE which defines the maximum size a single memory access may have (used for amongst other setting the size of the buffers), USE_DATA which will enable the infrastructure for fetching and sending the accessed data in the events, MULTITHREADED which will enable the still incomplete multithreaded code, USE_STATES which will enable the fast forward, warm up and simulation state machine and DTRACE which will make pint output some debugging information. Given the impact these features can have we decided to allow the user disable them at compile time. Also some of these features can be disabled at run time although they will still have some impact on the execution, in particular, USE_STATES will still cause the slowdowns of the conditionals introduced before the instrumentation calls to handle the state machine and USE_DATA will make the event size, and thus the queues larger. A final option that can be disabled is mapping tracing after a context 23
Chapter 6 Conclusions The project development has taken a long time given the research components it had yet, its development helped us have a good insight on how to improve simulators speed. Also, given the promising results obtained with the benchmarks (worst case of 804 when simulating, best case of 199 with a mean of 360,5 and an average of 419) run during the development and testing of this fairly limited version we think that ideas like simulation segmentation and instrumentation based simulation on independent process may help to the development of faster and more powerful simulators and will continue with its development. We expect to see in the future heavily multithreaded simulators where each processor has its own group of threads each handling the different stages independently in order to speed up execution times on multiprocessor machines. We also expect to see in the future more simulators used the process based separation between the data collection routines responsible of the execution of the program and the simulation itself in order to allow for the usage of higher level languages with less restrictions whilst still providing high performance 30
Chapter 6. Conclusions 6.1. Improvements for next release and native execution of the simulated programs. 6.1 Improvements for next release In the next release we intend to have a fully parallel instrumentation framework, we will also reimplement the trace simulator as a full gem5 CPU so it can have proper TLB handling and can be extended internally with more complex models. Finally we will change the queuing system so the simulator knows how many events will be generated by the instruction being executed before these events are handled down. We will also interconnect Multi2sim with gem5 in order to prove the powerfulness of Mead. Once we release the next version we intend to publish a paper on a publication on this topic. 31
Appendix A User manual A.1 Building In order to build the sources it suffices with running the make command on the sources directory. A.2 Pint Running the Pint pintool is quite easy and for that it is enough to run: ./pin -t source/tools/SimpleExamples/obj-intel64/pinatrace.so – command arguments Options can be set by setting the desired switches between pinatrace.so and the – Currently the following options are available: -f number : adds the set number of instruction to be run in the fast forward state (used many times it will set more instruction counts to be run the next time we go back to said state) -w number : adds the set number of instruction to be run in the warm up 32
Appendix A. User manual A.3. Schnapps state (used many times it will set more instruction counts to be run the next time we go back to said state) -s number : adds the set number of instruction to be run in the simulation state (used many times it will set more instruction counts to be run the next time we go back to said state) -syscallmap {0,1} : disables, if 0, or enables, if 1, the checking of process mappings after returning from a syscall -ctxchangemap {0,1} : disables, if 0, or enables, if 1, the checking of process mappings after a context change -values {0,1}: disables, if 0, or enables, if 1, the copying of data along with the memory events A.3 Schnapps For running Schnapps just run ./consumer A.4 Gin5 Gin5 requires a python file setting the system to be emulated. An example of such system can be found in the pintrace.py file. Once you have set up your system on a python file you just need to run the gem5.fast binary followed by the file containing the system being defined. Scripts to set up systems may take arguments from the command line if introduced after the script file. Our example file does not make use of this feature but others may. 33
Appendix B Relevant source code pinatrace.h 1#ifndef PINATRACE_H 2#define PINATRACE_H 3#include <l i n u x / f u t e x . h> 4#include <s y s / i p c . h> 5#include <s y s /sem . h> 6#include <s y s /shm . h> 7#include <s y s / s y s c a l l . h> 8#include <s y s / ty pes . h> 9#include <u n i s t d . h> 10 11 #include <c s i g n a l > 12 #include <c s t d i o > 13 #include <c s t d l i b > 14 #include <c s t r i n g > 15 #include <new> 16 17 #define l i k e l y ( x ) __builtin_expect ( ! ! ( x ) , 1) 18 #define u n l i k e l y ( x ) __builtin_expect ( ! ! ( x ) , 0 ) 19 20 // Compile ti me c o n f i g u r a t i o n 21 #define PADSIZE 64 // 64 b y t e l i n e s i z e as p er i n t e l s p ec s 22 #define MAXMEMSIZE 32 // 256 b i t s as p er AVX, n eed s t o be i n c r e a s e d on t h e f u t u r e 23 //#d e f i n e USE_DATA 1 // Add c o d e p a t h s t o o b t a i n t h e d a t a on t h e memory a c c e s s e s 24 //#d e f i n e MULTITHREADED 1 // Add c od e p at h s t o a l l o w f o r m u l t i t h r e a d e d a p p l i c a t i o n s 25 //#d e f i n e USE_STATES 1 // Use a s t a t e machine t o a l l o w f o r f a s t f o r w a r d and warm up s t a t e s 26 //#d e f i n e DTRACE 1 // Add d e b u g i n g o u t p u t 27 //#d e f i n e USE_YIELD 1 // Use y i e l d ( ) or t h e p i n e q u i v a l e n t when w a i t i n g f o r t h e o t h e r t h r e a d 28 #define QSIZE_ 1024 // D e f a u l t FIFO que ue s i z e 29 30 #ifdef USE_YIELD 31 #ifdef PIN_H 32 #define YIELD PIN_Yield 33 #else 34 #include <sched . h> 35 #define YIELD sched_yield 36 #endif 37 #else 38 #define YIELD( ) 39 #endif 40 41 #define c a c h e a l i g n e d __attribute__ ( ( a l i g n e d (PADSIZE ) ) ) 42 43 44 #ifndef PIN_H 45 typedef void VOID; 46 typedef u_int32_t UINT32 ; 47 typedef u_int8_t UINT8 ; 48 #endif 49 50 enum DataType { INVALDATA, STARTTH, ACCMEM } ; 34
Appendix B. Relevant source code 51 enum AccessType {ACCEXEC, ACCREAD, ACCWRITE, ACCPREFETCH } ; 52 53 54 // #d e f i n e PAD( n) ( ( ( ( n ) + ( PADSIZE −1 ) ) / PADSIZE ) ∗PADSIZE) 55 56 #include <c a s s e r t > 57 58 #ifdef DTRACE 59 #define d c p r i n t f ( c , . . . ) i f ( c ) f p r i n t f ( st d e r r , __VA_ARGS__) 60 #define d p r i n t f ( . . . ) f p r i n t f ( s t de r r , __VA_ARGS__) 61 #define dcputs ( c , a ) i f ( c ) f pu t s ( ( a ) , s t d e r r ) 62 #define dputs ( a ) f p ut s (( a ) , s t d e r r ) 63 #else 64 #define dcprintf (...) 65 #define d p r i n t f ( . . . ) 66 #define dcputs ( c , a ) 67 #define dputs ( a ) 68 #endif 69 70 //TODO: u se a l i g n m e n t s i n s t e a d o f p a d d i ng s 71 //TODO: u se o t h e r padded s t r u c t f o r t he da t a from re ad t o w r i t e 72 class MemAccess { 73 private : 74 //We ar e not g o i n g t o u se d e r i v a t e c l a s s e s her e f o r e f f i c i e n c y 75 AccessType type ; 76 VOID ∗ea ; // E f f e c t i v e a dd r es s o f t h e a c ce s s 77 #i f d e f USE_DATA 78 char data [MAXMEMSIZE ] ; // C ont ai ns e i t h e r t h e d a ta e x e c u t e d / re ad o r t h e d a ta c o n t ai n e d b e f o r e w r i t i n g 79 char wdata [MAXMEMSIZE] ; // Thi s i s v a l i d o n ly when t h e d at a a c c e s s i s a w r i t e c o n t a i n s t h e w r i t t e n d a ta 80 #endif 81 UINT32 s i z e ; // S i z e o f t h e a c c e s s 82 #i f d e f MULTITHREADED 83 UINT32 t i d ; // The ID o f t h e t h r e a d g e n e r a t i n g t h e a c c e s s 84 #endif 85 #i f d e f USE_DATA 86 i n l i n e void se tData ( ) { 87 a s s e r t ( PIN_SafeCopy ( data , ea , s i z e ) == s i z e ) ; 88 } 89 i n l i n e void copyData ( const MemAccess &ma) { 90 memcpy( data , ma. data , s i z e ) ; 91 i f ( type == ACCWRITE ) 92 memcpy( wdata , ma . wdata , s i z e ) ; 93 } 94 #endif 95 public : 96 #i f d e f USE_DATA 97 i n l i n e void setWdata ( ) { 98 a s s e r t ( PIN_SafeCopy ( wdata , ea , s i z e ) == s i z e ) ; 99 } 100 #endif 101 i n l i n e void MemAccessSet ( AccessType type , VOID ∗ea , UINT32 s i z e 102 #i f d e f MULTITHREADED 103 , UINT32 t i d 104 #endif 105 ) { 106 this−>type = type ; 107 this−>ea = ea ; 108 this−>s i z e = s i z e ; 109 #i f d e f MULTITHREADED 110 this−>t i d = t i d ; 111 #endif 112 #i f d e f USE_DATA 113 i f ( type != ACCPREFETCH) { 114 setData ( ) ; 115 } 116 #endif 117 } 118 i n l i n e void MemAccessSet (const MemAccess &ma) { 119 type = ma. type ; 120 ea = ma. ea ; 121 s i z e = ma. s i z e ; 122 #i f d e f MULTITHREADED 123 t i d = ma. t i d ; 124 #endif 125 #i f d e f USE_DATA 126 i f ( type != ACCPREFETCH) { 127 copyData (ma ) ; 128 } 129 #endif 130 } 131 void show ( FILE ∗f ) ; // R e q u i r e s t h e C LOCK 132 inline AccessType getType ( ) const {return type ; } 133 i n l i n e void ∗getEA ( ) const {return ea ; } 35
Appendix B. Relevant source code 134 #i f d e f USE_DATA 135 i n l i n e const void ∗getData ( ) const {return d at a ; } 136 i n l i n e const void ∗getWData () const {return wdata ; } 137 #endif 138 inline UINT32 g e t S i z e ( ) const {return s i z e ; } 139 #i f d e f MULTITHREADED 140 inline UINT32 g etTid ( ) const {return t i d ; } 141 #endif 142 } cachealigned ; 143 144 union SimDataU { 145 c l a s s MemAccess ma; 146 } ; 147 148 class SimData { 149 private : 150 DataType type ; 151 SimDataU data ; 152 public : 153 SimData ( ) : t ype (INVALDATA) { 154 } 155 SimData ( DataType _type ) : type ( _type ) { 156 } 157 inline DataType getType ( ) const { 158 return type ; 159 } 160 i n l i n e void setType ( DataType _type ) { 161 type = _type ; 162 } 163 inline MemAccess & getMa ( ) { 164 type = ACCMEM; 165 return data .ma ; 166 } 167 i n l i n e const MemAccess & getCMa ( ) const { 168 a s s e r t ( type == ACCMEM) ; 169 return data .ma ; 170 } 171 } ; 172 173 enum InstEventType { 174 ADDMAPPING, //A mapping was added d urin g t h e l a s t c o n t e x t cha nge / s y s c a l l 175 REMOVEMAPPING, //A mapping was removed d ur ing t h e l a s t c o n t e x t chang e / s y s c a l l 176 ADDTHREAD, //A new e x e c u t i o n t h r e ad ha s b een spawned 177 REMOVETHREAD, //An e x e c u t i o n t h r ea d h as c ea sed e x i s t i n g 178 } ; 179 180 struct range { 181 unsigned long i n t b ; //begin 182 unsigned long i n t e ; // end 183 i n l i n e bool operator < ( const struct range &r ) const { 184 // There s ho ul d n ’ t be o v e r l a p p i n g r an ge s ( at l e a s t i n t h e or y ) ; 185 return b < r . b ; 186 } 187 } ; 188 189 union InstEventData { 190 range r ; 191 //TODO: h a n dl e pe r t h r e a d a c c e s s qu eue c r e a t i o n h e r e 192 } ; 193 194 195 class In stEve nt { 196 private : 197 InstEventType type ; 198 InstEventData data ; 199 public : 200 inline In stE ven t ( ) { } 201 i n l i n e void S et I ns t Ev en t ( I ns tE ve nt Ty pe _type , r ange _r ) { 202 type = _type ; 203 data . r = _r ; 204 } 205 inline InstEventType getType ( ) const { 206 return type ; 207 } 208 inline rang e getRange ( ) { 209 a s s e r t ( type == ADDMAPPING | | type == REMOVEMAPPING) ; 210 return data . r ; 211 } 212 } cachealigned ; 213 // Th is i s a c l a s s i m p le m en t in g l o c k l e s s s i n g l e p r o d u c e r s s i n g l e c onsumer q u e ue s 214 // They a re v e r y u s e f u l f o r f a s t e f f i c i e n t IPC t h r o u g h s h a r e d memory t h ou g h you 215 // need t o e nsu r e t he s t r u c t u r e b ei n g que ued has a l l t h e n e ce s s a ry da ta i n s i d e 216 // i . e . d oesn ’ t u se s r e f e r e n c e s . 36
Appendix B. Relevant source code 217 // C u r r e n t l y we u se them f o r two p ur po us es , p a s s i n g e v e n t s r e l a t e d t o memory 218 // mappings and t h r e a d s b et wee n t h e i n s t r u m e n t a t i o n and t h e s i m u l a t o r and 219 // p a s s i n g around t h e memory a c c e s e s o f ea ch t h r e ad . 220 221 //We o n l y u se QSIZE −1 t hu s t h e r e i s alw ay s one e le me nt f r e e f o r p r o c e s s i n g b e f o r e q u euei n g . 222 #define NEXTQELEM( v ) ( ( ( v ) + 1) % QSIZE) 223 224 enum shmstatus {NONE = 0 , // I n i t i a l s t a t e 225 CLIENT_ACK=1, //The c l i e n t c on f irm s r e c e p t i o n o f p r e v i o u s s t a t e 226 SERVER_STARTED=2, //The s e r v e r ha s j u s t s t a r t e d 227 SERVER_DIED=3, //The s e r v e r has d ie d 228 // T his one s r e f e r t o t h e n e xt i n s t r u c t i o n pu shed t o t he qu eue ( so t h ey i n c l u d e up u n t i l t he ACCEXEC a f t e r t h a t ) 229 SERVER_SIM_START=4, //We ar e g o in g t o jump i n t o s i m u l a t i o n r e s e t s t a t s 230 SERVER_SIM_END=5 //We have ended s i m u l a t i o n r e s e t s t a t s 231 } ; 232 233 234 // A l o c k l e s s s i n g l e p r o du c e r s i n g l e consumer queue , w i t h more th an 1 you w i l l need l o c k s 235 template <class T, int QSIZE=QSIZE_> class SHMQ { 236 T queue [ QSIZE ] c a c h e a l i g n e d ; 237 v o l a t i l e s ig_ ato mic_ t qhead c a c h e a l i g n e d ; 238 v o l a t i l e s ig_ ato mic_ t q t a i l c a c h e a l i g n e d ; 239 // Ele men ts a re i n s e r t e d on t h e head and removed from t h e t a i l l i k e a snake . 240 v o l a t i l e s ig_ ato mic_ t c o n t r o l c a c h e a l i g n e d ; 241 public : 242 inline SHMQ ( ) : qhead ( 0 ) , q t a i l ( 0 ) { 243 } 244 inline T & get head ( ) { return queue [ qhead ] ; } 245 inline T & g e t t a i l ( ) { 246 a s s e r t ( ! empty ( ) ) ; 247 return queue [ q t a i l ] ; 248 } 249 i n l i n e bool f u l l ( ) {return NEXTQELEM( qhead ) == q t a i l ; } 250 i n l i n e bool empty ( ) {return q t a i l == qhead ; } 251 // Wait f o r t he qu eue no t t o be f u l l 252 i n l i n e void w a i t _ f u l l ( ) { 253 while ( u n l i k e l y ( f u l l ( ) ) ) YIELD ( ) ; 254 } 255 // Wait f o r t he qu eue not t o be empty ( I f t h e s e r v e r d i e s i t w i l l n ev er be ) 256 i n l i n e bool wait_empty_cond () { 257 return ( empty ( ) && c o n t r o l == CLIENT_ACK) ; 258 } 259 i n l i n e void wait_empty ( ) { 260 while ( u n l i k e l y ( wait_empty_cond ( ) ) ) YIELD ( ) ; 261 } 262 i n l i n e void wait_not_empty ( ) { 263 while ( u n l i k e l y ( ! empty ( ) ) ) YIELD ( ) ; 264 } 265 i n l i n e void push ( ) { 266 a s s e r t ( ! f u l l ( ) ) ; 267 qhead = NEXTQELEM( qhead ) ; 268 } 269 // Wait i f n e c e s s a r y t he n p ush 270 i n l i n e void wait_push ( ) { 271 w a i t _ f u l l ( ) ; 272 push ( ) ; 273 } 274 i n l i n e void pop ( ) { 275 a s s e r t ( ! empty ( ) ) ; 276 q t a i l = NEXTQELEM( q t a i l ) ; 277 } 278 i n l i n e enum shmstatu s r e c e i v e _ c o n t r o l ( ) { 279 i f ( c o n t r o l == CLIENT_ACK) return NONE; 280 return (enum shmstatu s ) c o n t r o l ; 281 } 282 i n l i n e void ack_control () { 283 while ( u n l i k e l y ( c o n t r o l == CLIENT_ACK) ) YIELD ( ) ; 284 c o n t r o l = CLIENT_ACK; 285 } 286 i n l i n e void send_control (enum shmstatus st ) { 287 a s s e r t ( s t != CLIENT_ACK ) ; // For t h i s we s h o u l d u se a c k _ c o n t r o l i n s t e a d 288 c o n t r o l = s t ; 289 // Wait f o r t h e ACK 290 while ( u n l i k e l y ( c o n t r o l != CLIENT_ACK) ) YIELD ( ) ; 291 } 292 i n l i n e void w ai t_ st ar t ( ) { 293 sig_atomic_t c on trol_ ; 294 while ( ( c on tr ol _ = c o n t r o l ) != SERVER_STARTED) s y s c a l l ( SYS_futex , &c o n t r o l ,FUTEX_WAIT, con trol_ , NULL,NULL , 0 ) ; 295 c o n t r o l = CLIENT_ACK; 296 s y s c a l l ( SYS_futex , &c o n t r o l ,FUTEX_WAKE, 1 ,NULL, NULL, 0 ) ; 297 } 298 i n l i n e void t e l l _ s t a r t ( ) { 299 sig_atomic_t c on trol_ ; 37
Appendix B. Relevant source code 300 c o n t r o l = SERVER_STARTED; 301 s y s c a l l ( SYS_futex , &c o n t r o l ,FUTEX_WAKE, 1 ,NULL, NULL, 0 ) ; 302 // Wait f o r t h e ACK 303 while ( ( con tro l_ = c o n t r o l ) != CLIENT_ACK) s y s c a l l ( SYS_futex , &c o nt r ol ,FUTEX_WAIT,SERVER_STARTED,NULL, NULL, 0 ) ; 304 } 305 } ; 306 307 typedef SHMQ<SimData> SimDataq ; 308 typedef SHMQ<InstEvent > InstEventq ; 309 310 SimDataq ∗s e r v e r _ i n i t 2 ( ) ; 311 SimDataq ∗c l i e n t _ i n i t 2 ( ) ; 312 void s e r v e r _ f i n i 2 ( SimDataq ∗q ) ; 313 void c l i e n t _ f i n i 2 ( SimDataq ∗q ) ; 314 315 //TODO: F ix t h e c a s e where t h e c l i e n t i s t h e one d o i ng t h e f i n a l i z a t i o n 316 317 //TODO: a c c e s s q ue ue s s h o u l d b e c r e a t e d d y n a m i c a l l y and p a s se d t h r o u gh t h e e v e n t qu eue 318 SimDataq ∗get_q2( in t &shmid ) { 319 SimDataq ∗q ; 320 i f ( ( shmid = shmget ( 26 84 , sizeof( SimDataq ) , IPC_CREAT | 066 6) ) < 0) { 321 p e r r o r ( " shmget " ) ; 322 e x i t ( 1 ) ; 323 } 324 void ∗shm ; 325 i f ( ( shm = shmat ( shmid , NULL, 0 ) ) == ( void ∗)−1) { 326 p e r r o r ( " shmat " ) ; 327 e x i t ( 1 ) ; 328 } 329 q = static_cast<SimDataq∗>(shm ) ; 330 return q ; 331 } 332 333 334 SimDataq ∗s e r v e r _ i n i t 2 ( ) { 335 int shmid ; 336 SimDataq ∗q = get_q2 ( shmid ) ; 337 new ( q ) SimDataq ( ) ; //We u se a p l a ce m e nt new s o we h av e t h e SimDataq i n t h e s h a r e d memory 338 q−>t e l l _ s t a r t ( ) ; 339 return q ; 340 } 341 342 SimDataq ∗c l i e n t _ i n i t 2 ( ) { 343 int shmid ; 344 SimDataq ∗q = get_q2 ( shmid ) ; 345 q−>w ait _st art ( ) ; 346 // S i n c e we a re c o n n e c t e d we t e l l t h e OS t h e s eg em en t can b e d e l e t e d 347 i f ( shmctl ( shmid , IPC_RMID,NULL) < 0) 348 p e r r o r ( " shmctl " ) ; 349 return q ; 350 } 351 352 void s e r v e r _ f i n i 2 ( SimDataq ∗q ) { 353 q−>send _c on trol (SERVER_DIED ) ; 354 } 355 356 void c l i e n t _ f i n i 2 ( SimDataq ∗q ) { 357 q−>ack_control (); 358 q−>~SimDataq ( ) ; 359 } 360 361 //TODO: w i t h p r op pe r t e m p l a t e u sa ge t h i s c o u l d g e t p r e t t i e r 362 Ins tEventq ∗server_init (); 363 Ins tEventq ∗c l i e n t _ i n i t ( ) ; 364 void s e r v e r _ f i n i ( I nst Eve ntq ∗q ) ; 365 void c l i e n t _ f i n i ( In stE ven tq ∗q ) ; 366 367 Ins tEventq ∗get_q ( i nt &shmid ) { 368 Ins tEventq ∗q ; 369 i f ( ( shmid = shmget ( 26 87 , sizeof( Inst Eve ntq ) , IPC_CREAT | 06 6 6) ) < 0) { 370 p e r r o r ( " shmget " ) ; 371 e x i t ( 1 ) ; 372 } 373 void ∗shm ; 374 i f ( ( shm = shmat ( shmid , NULL, 0 ) ) == ( void ∗)−1) { 375 p e r r o r ( " shmat " ) ; 376 e x i t ( 1 ) ; 377 } 378 q = static_cast<I nst Eve nt q ∗>(shm ) ; 379 return q ; 380 } 381 382 38
Appendix B. Relevant source code 383 Ins tEventq ∗server_init () { 384 int shmid ; 385 Ins tEventq ∗q = get_q ( shmid ) ; 386 new ( q ) I nstE vent q ( ) ; //We us e a p la ce me nt new s o we ha ve t h e I n s t E v e n t q i n t h e s ha r ed memory 387 q−>t e l l _ s t a r t ( ) ; 388 return q ; 389 } 390 391 Ins tEventq ∗c l i e n t _ i n i t ( ) { 392 int shmid ; 393 Ins tEventq ∗q = get_q ( shmid ) ; 394 q−>w ait _st art ( ) ; 395 // S i nce we ar e c on nec te d we t e l l t h e OS t h e s ege ment can b e d e l e t e d 396 i f ( shmctl ( shmid , IPC_RMID,NULL) < 0) 397 p e r r o r ( " shmctl " ) ; 398 return q ; 399 } 400 401 void s e r v e r _ f i n i ( I nst Eve ntq ∗q ) { 402 q−>send _c on trol (SERVER_DIED ) ; 403 } 404 405 void c l i e n t _ f i n i ( In stE ven tq ∗q ) { 406 q−>ack_control (); 407 q−>~I nst Eve ntq ( ) ; 408 } 409 410 411 #endif 39
Appendix B. Relevant source code 498 IARG_MEMORYREAD_SIZE, 499 #i f d e f MULTITHREADED 500 IARG_THREAD_ID, 501 #endif 502 IARG_END) ; 503 } 504 505 // i ns tr u me n ts s t o r e s u s in g a p r e d i c a t e d c a l l , i . e . 506 // t h e c a l l hap pens i f f t h e s t o r e w i l l be a c t u a l l y e x e c u t e d 507 i f (INS_IsMemoryWrite( ins )) 508 { 509 #ifdef USE_STATES 510 I N S _ I n s e r t I f P r e d i c a t e d C a l l ( i ns , IPOINT_BEFORE, (AFUNPTR) Instrument , IARG_FAST_ANALYSIS_CALL, 511 #i f d e f MULTITHREADED 512 IARG_THREAD_ID, 513 #endif 514 IARG_END) ; 515 INS_InsertThenPredicatedCall 516 #else 517 INS_InsertPredicatedCall 518 #endif 519 ( i ns , IPOINT_BEFORE, (AFUNPTR) RecordMemPreWrite , IARG_FAST_ANALYSIS_CALL, 520 IARG_MEMORYWRITE_EA, 521 IARG_MEMORYWRITE_SIZE, 522 #i f d e f MULTITHREADED 523 IARG_THREAD_ID, 524 #endif 525 IARG_END) ; 526 #ifdef USE_DATA 527 #ifdef USE_STATES 528 I N S _ I n s e r t I f P r e d i c a t e d C a l l ( i ns , IPOINT_BEFORE, (AFUNPTR) Instrument , IARG_FAST_ANALYSIS_CALL, 529 #i f d e f MULTITHREADED 530 IARG_THREAD_ID, 531 #endif 532 IARG_END) ; 533 INS_InsertThenPredicatedCall 534 #else 535 INS_InsertPredicatedCall 536 #endif 537 ( in s , 538 ( ! INS_HasFallThrough ( i n s )?IPOINT_TAKEN_BRANCH: IPOINT_AFTER) , 539 (AFUNPTR) RecordMemWrite , IARG_FAST_ANALYSIS_CALL, 540 #i f d e f MULTITHREADED 541 IARG_THREAD_ID, 542 #endif 543 IARG_END) ; 544 #endif 545 } 546 // } 547 } 548 549 // Multithread stuff : 550 551 552 #ifdef MULTITHREADED 553 s t a t i c VOID ThreadStart (THREADID t id , CONTEXT ∗c txt , INT32 f l a g s , VOID ∗v ) 554 { 555 MemAccess ∗ma = new MemAccess ( ) ; 556 PIN_SetThreadData ( wMemAccess , ma, t i d ) ; 557 } 558 559 560 s t a t i c VOID ThreadFini (THREADID t id , const CONTEXT ∗c txt , INT32 code , VOID ∗v ) 561 { 562 MemAccess ∗ma = static_cast<MemAccess ∗>(PIN_GetThreadData ( wMemAccess , t i d ) ) ; 563 delete ma; 564 } 565 #endif 566 567 //TODO: maybe i n t e g r a t e t h i s i n t o t h e que ue c l a s s and t h e s o c k e t p er t h r e a d p r o t o c o l 568 // s t a t i c b o o l end i ng = f a l s e ; 569 // s t a t i c THREADID p r o c e s s o r ; 570 // s t a t i c PIN_THREAD_UID p r o c e s s o r u i d ; 571 // 572 // s t a t i c VOID ProcessQueue (VOID ∗n o t h i n g ) { 573 // THREADID t i d = PIN_ThreadId ( ) ; 574 // w h i l e ( ! en di ng | | ! q−>empty ( ) ) { 575 // GetL oc k(& c _l ock , t i d ) ; 576 // w h i l e ( ! q−>empty ( ) ) { 577 // q−>gettail (); 578 // q−>pop ( ) ; 579 // } 580 // R el ea se L oc k (& c _l oc k ) ; 46
Appendix B. Relevant source code 581 // // Let o t h e r s f i l l t h e queue 582 // YIELD ( ) ; 583 // } 584 // } 585 586 s t a t i c VOID F i n i ( INT32 code , VOID ∗v ) 587 { 588 // e nd in g = t r u e ; 589 // PIN_WaitForThreadTermination ( p r o c e s s o r u i d , PIN_INFINITE_TIMEOUT, NULL ) ; 590 i f ( KnobSyscallMap | | KnobCtxChangeMap ) 591 f c l o s e ( mout ) ; 592 server_fini2(q); 593 server_fini(iq ); 594 } 595 596 int main( i nt argc , char ∗argv [ ] ) 597 { 598 i f ( PIN_Init ( argc , argv ) ) 599 { 600 return Usage ( ) ; 601 } 602 603 #ifdef USE_STATES 604 i f ( ! ( Knobf . NumberOfValues ( ) == Knobw . NumberOfValues ( ) && Knobf . NumberOfValues()==Knobs . NumberOfValues ( ) ) ) 605 { 606 f p u t s ( " The␣number␣ o f ␣ o c c u r r e n c e s ␣ o f ␣−f ␣−h␣and␣−s ␣must␣be␣ the ␣same . " , s t d e r r ) ; 607 return Usage ( ) ; 608 } 609 #endif 610 611 iq=server_init (); 612 q=s e r v e r _ i n i t 2 ( ) ; 613 q−>gethead ( ) . setType (STARTTH) ; //TODO move t o t he t h r ea d s t a r t c a l l b a c k s 614 q−>wait_push ( ) ; 615 616 #ifdef USE_STATES 617 i f ( Knobf . NumberOfValues ( ) >= 1 ) 618 nex t St a te ( ) ; 619 // T his one i s done due t o t h e way i n s t r u m e n t a t i o n wor ks 620 inscount++; 621 // Send t h e simu s t a r t command i f n e c e s s a r y 622 i f ( s t a t e == SIMULATION) 623 #endif 624 q−>send _c on trol (SERVER_SIM_START) ; 625 INS_AddInstrumentFunction( Instruction , 0); 626 PIN_AddFiniUnlockedFunction( Fini , 0); 627 628 // Open t h e o u t p u t f i l e 629 i f ( KnobSyscallMap | | KnobCtxChangeMap ) 630 mout = fopen ( " maptrace . t xt " ,"w" ) ; 631 632 633 // Monitor s y s c a l l s and s o f o r mapping c ha n ge s 634 i f ( KnobSyscallMap ) 635 PIN_AddSyscallExitFunction(parsemaps1 , NULL); 636 // Monitor a l s o a f t e r c o n t e x t c ha nge s s i n c e i f we a re p t ra ce d mappings may have c hange d 637 i f (KnobCtxChangeMap) 638 PIN_AddContextChangeFunction ( parsemaps2 , NULL ) ; 639 // A lth o ug h p i n u se s c od e c a c h e s i t h i d e s t h i s d e t a i l s from t h e i n s t r u m e nt a t i o n code so our i n s t r u c t i o n s c ach es don ’ t break . 640 // T his means t h e i n s t r u c t i o n a d d r e s s e s we g e t a re mapped t o t h e m ap pi ng s c o r r e s p o n d i n g t o t h e l i b r a r i e s and n ot t o t h e JIT 641 // c ode so we don ’ t have t o worry a bou t ch a ng es t o t h e s e mappings , but , s i n c e we s t i l l can ’ t d i s c e r n them from a p p l i c a t i o n 642 // mappings we s t i l l ha ve t o r e s e r v e spa ce f o r them i n t h e s i m u l a t o r s pac e . Th is a l s o means we ’ l l be h a vi n g some movement 643 // i n t he map spa c e a lm ost a l w ay s u n t i l PIN p r o v i d e s an a p i t o d i s c e r n p in / t o o l mappings from a p p l i c a t i o n one s . 644 645 // Thread C a l l b a c k s 646 I n i t L o c k (& c_lock ) ; 647 648 #i f d e f MULTITHREADED 649 I n i t L o c k (&h_lock ) ; 650 wMemAccess = PIN_CreateThreadDataKey ( 0 ) ; 651 PIN_AddThreadStartFunction ( ThreadStart , 0 ) ; 652 PIN_AddThreadFiniFunction ( ThreadFini , 0 ) ; 653 #endif 654 655 // S t a r t qu eue p r o c e s s o r t h r e a d 656 // p r o c e s s or = PIN_SpawnInternalThread ( ProcessQueue , NULL, 0 , &p r o c e s s o r u i d ) ; 657 // i f ( p r o c e s s o r == INVALID_THREADID) r e t u r n 1 ; 658 // I n i t i a l map l o a d i n g 659 i f ( KnobSyscallMap | | KnobCtxChangeMap ) 660 parsemaps3 (); 661 PIN_StartProgram ( ) ; 662 47
Appendix B. Relevant source code 663 return 0 ; 664 } 48
Appendix B. Relevant source code consumer.cpp 1#include <pinatrace .h> 2#include <cstdint> 3#include <c i n t t y p e s > 4 5int main ( ) { 6Ins tEventq ∗i q ; 7i q = c l i e n t _ i n i t ( ) ; 8SimDataq ∗q ; 9q= c l i e n t _ i n i t 2 ( ) ; 10 uint64_ t n in s = 0 ; 11 uint64_ t n in s2 = 0 ; 12 uint64_ t nr ea = 0 ; 13 uint64_ t nr ea2 = 0 ; 14 uint64_ t nwri = 0 ; 15 uint64_ t nwri2 = 0 ; 16 uint64_ t npre = 0 ; 17 uint64_ t npre2 = 0 ; 18 uint64_t s i n s = 0 ; 19 uint64_t s i n s 2 = 0 ; 20 uint6 4_t s r e a = 0 ; 21 uin t64_t s r e a 2 = 0 ; 22 uint64_ t s wr i = 0 ; 23 uint64_ t s wr i2 = 0 ; 24 uint64_ t s pr e = 0 ; 25 uint64_ t s pr e2 = 0 ; 26 uint64_ t mark = 0 ; 27 uint64_ t mark2 = 0 ; 28 bool simulating = false ; 29 while ( q−>r e c e i v e _ c o n t r o l ( ) != SERVER_DIED) { 30 while ( ! q−>empty ( ) ) { 31 a s s e r t (q−>g e t t a i l ( ) . getType ( ) != INVALDATA) ; 32 i f ( q−>g e t t a i l ( ) . getType ( ) == ACCMEM) { 33 const MemAccess &ma = q−>g e t t a i l ( ) . getCMa ( ) ; 34 mark ^= ( uint64_t ) ma. getEA ( ) ; 35 mark2 ^= ( uint64_t ) ma. getEA ( ) ; 36 switch(ma. getType ( ) ) { 37 case ACCEXEC: 38 nin s ++; 39 ni n s2 ++; 40 s i n s += ma . g e t S i z e ( ) ; 41 s i n s 2 += ma . g e t S i z e ( ) ; 42 break ; 43 case ACCREAD: 44 nrea ++; 45 nrea2++; 46 s r ea += ma . g e t S i z e ( ) ; 47 s re a 2 += ma . g e t S i z e ( ) ; 48 break ; 49 case ACCWRITE: 50 nwri ++; 51 nwri2++; 52 s wr i += ma . g e t S i z e ( ) ; 53 sw ri 2 += ma . g e t S i z e ( ) ; 54 break ; 55 case ACCPREFETCH: 56 npre ++; 57 npre2++; 58 s pr e += ma . g e t S i z e ( ) ; 59 sp re 2 += ma . g e t S i z e ( ) ; 60 break ; 61 default : 62 puts ( " Unexpected ␣ a c c e s s ␣ type ! " ) ; 63 } 64 } 65 q−>pop ( ) ; 66 } 67 // Have we j u s t em ptied t he b u f f e r or has an e v en t happened ? 68 while ( ! iq−>empty ( ) ) { 69 i f ( iq−>g e t t a i l ( ) . getType ( ) == REMOVEMAPPING){ 70 range r=iq−>g e t t a i l ( ) . getRange ( ) ; 71 printf("−␣%lx−%l x \n " , r . b , r . e ) ; 72 }e l s e i f ( iq−>g e t t a i l ( ) . getType ( ) == ADDMAPPING){ 73 range r=iq−>g e t t a i l ( ) . getRange ( ) ; 74 printf("+␣%lx−%l x \n " , r . b , r . e ) ; 75 } 76 iq−>pop ( ) ; 77 } 78 switch( q−>r e c e i v e _ c o n t r o l ( ) ) { 79 case SERVER_DIED: 80 i f ( ! s i m u l a t i n g ) 81 break ; 82 case SERVER_SIM_END: 49
Appendix B. Relevant source code 83 simulating = false ; 84 puts ( " S imu la ti o n ␣ s t a t i s t i c s : " ) ; 85 puts ( "Number␣ o f ␣ a c c e s s e s : " ) ; 86 printf(" ␣␣ i n s t r u c t i o n s : ␣%" PRIu64" \n " , ni n s ) ; 87 printf(" ␣␣ re a ds ␣␣␣␣␣␣␣ : ␣%" PRIu64" \n " , n rea ) ; 88 printf(" ␣ ␣ w r i t e s ␣␣ ␣␣␣␣ : ␣%" PRIu64" \n " , nwri ) ; 89 printf(" ␣␣ p r e f e t c h e s ␣␣ : ␣%" PRIu64 " \n " , npre ) ; 90 printf(" ␣␣ t o t a l ␣␣␣␣␣␣␣ : ␣%" PRIu64" \n " , n i ns+nrea+nwri+npre ) ; 91 puts ( " T otal ␣ a c c e s s e d ␣memory␣ by ␣ type ␣ ( b y te s ) : " ) ; 92 printf(" ␣␣ i n s t r u c t i o n s : ␣%" PRIu64" \n " , s i n s ) ; 93 printf(" ␣␣ re a ds ␣␣␣␣␣␣␣ : ␣%" PRIu64" \n " , s r e a ) ; 94 printf(" ␣ ␣ w r i t e s ␣␣ ␣␣␣␣ : ␣%" PRIu64" \n " , sw r i ) ; 95 printf(" ␣␣ p r e f e t c h e s ␣␣ : ␣%" PRIu64 " \n " , s p re ) ; 96 printf(" ␣␣ t o t a l ␣␣␣␣␣␣␣ : ␣%" PRIu64" \n " , s i n s+s r e a+s wri+s pr e ) ; 97 printf(" Execution ␣mark : ␣%" PRIx64 " \n " , mark ) ; 98 q−>ack_control (); 99 break ; 100 case SERVER_SIM_START: 101 n in s = 0 ; 102 nrea = 0 ; 103 nwri = 0 ; 104 npre = 0 ; 105 s i n s = 0 ; 106 s r e a = 0 ; 107 s wr i = 0 ; 108 s pr e = 0 ; 109 mark = 0 ; 110 simulating = true ; 111 q−>ack_control (); 112 break ; 113 } 114 // Wait f o r b u f f e r t o r e f i l l 115 while ( q−>wait_empty_cond ( ) && iq−>wait_empty_cond ( ) ) YIELD ( ) ; 116 } 117 puts ( " Total ␣ s t a t i s t i c s : " ) ; 118 puts ( "Number␣ o f ␣ a c c e s s e s : " ) ; 119 printf(" ␣␣ i n s t r u c t i o n s : ␣%" PRIu64" \n " , nins 2 ) ; 120 printf(" ␣␣ re a ds ␣␣␣␣␣␣␣ : ␣%" PRIu64" \n " , n rea2 ) ; 121 printf(" ␣ ␣ w r i t e s ␣␣ ␣␣␣␣ : ␣%" PRIu64" \n " , nwri2 ) ; 122 printf(" ␣␣ p r e f e t c h e s ␣␣ : ␣%" PRIu64 " \n " , npre2 ) ; 123 printf(" ␣␣ t o t a l ␣␣␣␣␣␣␣ : ␣%" PRIu64" \n " , n in s 2+nrea2+nwri2+npre2 ) ; 124 puts ( " T otal ␣ a c c e s s e d ␣memory␣ by ␣ type ␣ ( b y te s ) : " ) ; 125 printf(" ␣␣ i n s t r u c t i o n s : ␣%" PRIu64" \n " , s i n s 2 ) ; 126 printf(" ␣␣ re a ds ␣␣␣␣␣␣␣ : ␣%" PRIu64" \n " , s r e a 2 ) ; 127 printf(" ␣ ␣ w r i t e s ␣␣ ␣␣␣␣ : ␣%" PRIu64" \n " , swri 2 ) ; 128 printf(" ␣␣ p r e f e t c h e s ␣␣ : ␣%" PRIu64 " \n " , s pr e 2 ) ; 129 printf(" ␣␣ t o t a l ␣␣␣␣␣␣␣ : ␣%" PRIu64" \n " , s i n s 2+s r e a 2+s w r i2+s p re 2 ) ; 130 printf(" Execution ␣mark : ␣%" PRIx64 " \n " , mark2 ) ; 131 client_fini2(q); 132 c l i e n t _ f i n i ( iq ) ; 133 return 0 ; 134 } 50
Appendix B. Relevant source code mem_trace_reader.hh 1/∗ 2∗C o p y r i g h t ( c ) 2004 −2005 The R eg en ts o f The U n i v e r s i t y o f Michigan 3∗A l l r i g h t s r e s e r v e d . 4∗ 5∗R e d i s t r i b u t i o n and us e i n s ourc e and b in a r y forms , w it h or w i t h o u t 6∗m o d i f i c a t i o n , a re p e r m i t t e d p r o v i d e d t h a t t h e f o l l o w i n g c o n d i t i o n s a re 7∗met : r e d i s t r i b u t i o n s o f s o u r c e code must r e t a i n t he a bo ve c o p y r i g h t 8∗n ot i c e , t h i s l i s t o f c o n d i t i o n s and t h e f o l l o w i n g d i s c l a i m e r ; 9∗r e d i s t r i b u t i o n s in b i n a r y form must r e p r odu c e t he a bo ve c o p y r i g h t 10 ∗n ot i c e , t h i s l i s t o f c o n d i t i o n s and t h e f o l l o w i n g d i s c l a i m e r in t he 11 ∗do cu men ta ti on and/ or o t h e r m a t e r i a l s p r ov i d e d w i t h t h e d i s t r i b u t i o n ; 12 ∗n e i t h e r t he name o f t h e c o p y r i g h t h o l d e r s nor t he names o f i t s 13 ∗c o n t r i b u t o r s may be used t o e n do rs e or promote p ro d u ct s d e r i v e d from 14 ∗t h i s s o f t w a r e w i t h ou t s p e c i f i c p r i o r w r i t t e n p e rm i s si o n . 15 ∗ 16 ∗THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS 17 ∗"AS IS " AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT 18 ∗LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR 19 ∗A PARTICULAR PURPOSE ARE DISCLAIMED . IN NO EVENT SHALL THE COPYRIGHT 20 ∗OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT , INCIDENTAL , 21 ∗SPECIAL , EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT 22 ∗LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES ; LOSS OF USE, 23 ∗DATA, OR PROFITS ; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY 24 ∗THEORY OF LIABILITY , WHETHER IN CONTRACT, STRICT LIABILITY , OR TORT 25 ∗(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE 26 ∗OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. 27 ∗ 28 ∗Au tho rs : E r ik H a l l n o r 29 ∗/ 30 31 /∗ ∗ 32 ∗D e f i n i t i o n s f o r a pu re v i r t u a l i n t e r f a c e t o a memory t r a c e r ea de r . 33 ∗/ 34 35 #ifndef __MEM_TRACE_READER_HH__ 36 #define __MEM_TRACE_READER_HH__ 37 38 #include "mem/ pac ket . hh " 39 #include "mem/ r e q u e s t . hh " 40 #include " params/MemTraceReader . hh " 41 #include " sim / s im _obje ct . hh " 42 43 /∗ ∗ 44 ∗Thi s c l a s s c o n t a i n s t h e i n f o o f t h e t r a c e r e q u e s t and some u s e f u l methods t o 45 ∗s p l i t i t 46 ∗/ 47 class MemTraceRequest : public Fa stA l loc { 48 Addr _paddr ; 49 unsigned _s iz e ; 50 Reque st : : F l a gs _ f l a g s ; 51 Tick _time ; 52 int _asid ; 53 Addr _vaddr ; 54 int _contextId ; 55 int _threadId ; 56 Addr _pc ; 57 MemCmd _cmd ; 58 public : 59 MemTraceRequest () {} 60 MemTraceRequest ( Addr paddr , int s i z e , Request : : F lag s f l a g s , 61 MemCmd : : Command cmd) 62 : _paddr ( paddr ) , _ s i ze ( s i z e ) , _ f l a g s ( f l a g s ) , _time ( c urT ick ( ) ) , _cmd(cmd) 63 { } 64 65 MemTraceRequest ( Addr paddr , int s i z e , Request : : F lag s f l a g s , Tick time , 66 MemCmd : : Command cmd) 67 : _paddr ( paddr ) , _ s i ze ( s i z e ) , _ f l a g s ( f l a g s ) , _time ( time ) , _cmd(cmd) 68 { } 69 70 ~MemTraceRequest () {} // f o r F a s t A l l o c 71 72 /∗ ∗ 73 ∗Are we s c h e d u l e d t o run a l r e a d y 74 ∗/ 75 i n l i n e bool mustRun ( ) { 76 return _time <= curTic k ( ) ; 77 } 78 79 inline Tick time ( ) { 80 return _time ; 81 } 82 51
Appendix B. Relevant source code 83 i n l i n e bool isInstFetch () { 84 return _ f l a g s . i s S e t ( R eq ue st : : INST_FETCH ) ; 85 } 86 87 i n l i n e bool lastPacketSent () { 88 return _ si ze == 0 ; 89 } 90 91 /∗ ∗ 92 ∗Get t h e n e x t p a c k e t w i t h p r o p er b ounds f o r t h i s b l o c k s i z e 93 ∗W i l l r e t u r n NULL when done 94 ∗/ 95 PacketPtr getNextPkt ( i n t b s i z e , P ac ke t : : NodeID d est , MasterID mid ) { 96 i f (lastPacketSent ()) { 97 return NULL; 98 } 99 // Base a d d r e s s o f t h e b l o c k 100 Addr base = ( _paddr & ~( b s i z e −1 ) ) ; 101 // Cu rre nt b l o c k m axsi ze 102 int msize = b s i z e −( _paddr −base ) ; 103 //Minimum 104 i f ( msize > _ siz e ) msize = _s ize ; 105 // Ge nera te t t h e r e q u e s t and t h e p a c k e t 106 RequestPtr req = new Request ( _paddr , msize , _f lags , mid ) ; 107 PacketPtr pkt = new Packet ( req , _cmd, d e s t ) ; 108 pkt−>dataDynamicArray(new char [ msize ] ) ; 109 // C a l c u l a t e t he new b as e a dd r e s s and s i z e 110 _paddr += msize ; 111 _size −= msize ; 112 return pkt ; 113 } 114 } ; 115 116 typedef MemTraceRequest ∗MemTraceRequestPtr ; 117 118 119 /∗ ∗ 120 ∗Pure v i r t u a l b a se c l a s s f o r memory t r a c e r e a d e r s . 121 ∗/ 122 class MemTraceReader : public SimObject 123 { 124 public : 125 enum r e as o n {EOT,STAT_RESET,STAT_DUMP} ; 126 /∗ ∗ C o n s t r u c t t h i s MemoryTrace r e a d e r . ∗/ 127 MemTraceReader ( const MemTraceReaderParams ∗p ) : SimObject ( p ) {} 128 129 //TODO: r ed o doc p k t s h o u l d c o n t a i n time , r e q u e s t , command and d a ta . 130 /∗ ∗ 131 ∗Read t h e n e x t r e q u e s t from t h e t r a c e . Re tur ns t h e r e q u e s t i n t h e 132 ∗p r o v i d e d R e qu e st P tr and t h e c y c l e o f t h e r e q u e s t i n t h e r e t u r n v a l u e . 133 ∗@param r e q R etu rn t h e n e x t r e q u e s t from t h e t r a c e . 134 ∗@re turn The c y c l e o f t h e r e q u e s t , 0 i f none i n t r a c e . 135 ∗/ 136 virtual MemTraceRequestPtr getNextRequest (enum reason &reason) = 0; 137 138 } ; 139 140 #endif //__MEM_TRACE_READER_HH__ 52
Appendix B. Relevant source code pin_reader.hh 1/∗ 2∗C o p y r i g h t ( c ) 2004 −2005 The R eg en ts o f The U n i v e r s i t y o f Michigan 3∗A l l r i g h t s r e s e r v e d . 4∗ 5∗R e d i s t r i b u t i o n and us e i n s ourc e and b in a r y forms , w it h or w i t h o u t 6∗m o d i f i c a t i o n , a re p e r m i t t e d p r o v i d e d t h a t t h e f o l l o w i n g c o n d i t i o n s a re 7∗met : r e d i s t r i b u t i o n s o f s o u r c e code must r e t a i n t he a bo ve c o p y r i g h t 8∗n ot i c e , t h i s l i s t o f c o n d i t i o n s and t h e f o l l o w i n g d i s c l a i m e r ; 9∗r e d i s t r i b u t i o n s in b i n a r y form must r e p r odu c e t he a bo ve c o p y r i g h t 10 ∗n ot i c e , t h i s l i s t o f c o n d i t i o n s and t h e f o l l o w i n g d i s c l a i m e r in t he 11 ∗do cu men ta ti on and/ or o t h e r m a t e r i a l s p r ov i d e d w i t h t h e d i s t r i b u t i o n ; 12 ∗n e i t h e r t he name o f t h e c o p y r i g h t h o l d e r s nor t he names o f i t s 13 ∗c o n t r i b u t o r s may be used t o e n do rs e or promote p ro d u ct s d e r i v e d from 14 ∗t h i s s o f t w a r e w i t h ou t s p e c i f i c p r i o r w r i t t e n p e rm i s si o n . 15 ∗ 16 ∗THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS 17 ∗"AS IS " AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT 18 ∗LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR 19 ∗A PARTICULAR PURPOSE ARE DISCLAIMED . IN NO EVENT SHALL THE COPYRIGHT 20 ∗OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT , INCIDENTAL , 21 ∗SPECIAL , EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT 22 ∗LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES ; LOSS OF USE, 23 ∗DATA, OR PROFITS ; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY 24 ∗THEORY OF LIABILITY , WHETHER IN CONTRACT, STRICT LIABILITY , OR TORT 25 ∗(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE 26 ∗OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. 27 ∗ 28 ∗Au tho rs : E r ik H a l l n o r 29 ∗/ 30 31 /∗ ∗ 32 ∗@ f i l e 33 ∗D e f i n i t i o n o f a memory t r a c e r eader f o r a M5 memory t r a c e . 34 ∗/ 35 36 #ifndef __Pin_READER_HH__ 37 #define __Pin_READER_HH__ 38 39 #include " cpu/ t r a c e / r e a d e r / mem_trace_reader . hh " 40 #include " cpu/ t r a c e / r e a d e r / p in _a tr ac e . hh " 41 #include " params/ PinReader . hh " 42 43 /∗ ∗ 44 ∗A memory t r a c e r ea d er f o r a pin memory t r a c e . 45 ∗/ 46 class PinReader : public MemTraceReader 47 { 48 f ri e nd c l a s s DeleteQueuesCallback ; 49 /∗ ∗ The t r a c e . ∗/ 50 SimDataq ∗q ; 51 /∗ ∗ I n f o r m a t i o n on mapping c h a ng e s ∗/ 52 Ins tEventq ∗i q ; 53 bool simulating ; // Wether we a re i n s i m u l a t i o n s t a t e or no t 54 bool drop ; // S ho uld we d rop t h e n e x t e l em e nt ( h as i t bee n p r o c e s s e d ) 55 56 protected : 57 void removeQueues (); 58 public : 59 /∗ ∗ 60 ∗C o n s t r u c t an M5 memory t r a c e r e a d e r . 61 ∗/ 62 PinReader(const PinReaderParams ∗p ) ; 63 64 ~PinReader (); 65 66 67 //TODO: r ed o doc p k t s h o u l d c o n t a i n time , r e q u e s t , command and d a ta . 68 /∗ ∗ 69 ∗Read t h e n ex t r e q u e s t from t he t r a c e . Ret urns t h e r e q u e s t i n t he 70 ∗p r o v i d e d R e qu e st P tr and t h e c y c l e o f t h e r e q u e s t i n t h e r e t u r n v a l u e . 71 ∗@param r e q R et urn t h e n e x t r e q u e s t from t h e t r a c e . 72 ∗@re turn The c y c l e o f t h e r e q u e s t , 0 i f none i n t r a c e . 73 ∗/ 74 virtual MemTraceRequestPtr getNextRequest ( MemTraceReader : : re aso n &re aso n ) ; 75 } ; 76 77 #endif // __PIN_READER_HH__ 53
Appendix B. Relevant source code pin_reader.cc 1/∗ 2∗C o p y r i g h t ( c ) 2004 −2005 The R eg en ts o f The U n i v e r s i t y o f Michigan 3∗A l l r i g h t s r e s e r v e d . 4∗ 5∗R e d i s t r i b u t i o n and us e i n s ourc e and b in a r y forms , w it h or w i t h o u t 6∗m o d i f i c a t i o n , a re p e r m i t t e d p r o v i d e d t h a t t h e f o l l o w i n g c o n d i t i o n s a re 7∗met : r e d i s t r i b u t i o n s o f s o u r c e code must r e t a i n t he a bo ve c o p y r i g h t 8∗n ot i c e , t h i s l i s t o f c o n d i t i o n s and t h e f o l l o w i n g d i s c l a i m e r ; 9∗r e d i s t r i b u t i o n s in b i n a r y form must r e p r odu c e t he a bo ve c o p y r i g h t 10 ∗n ot i c e , t h i s l i s t o f c o n d i t i o n s and t h e f o l l o w i n g d i s c l a i m e r i n t he 11 ∗do cu men ta ti on and/ or o t h e r m a t e r i a l s p r ov i d e d w i t h t h e d i s t r i b u t i o n ; 12 ∗n e i t h e r t he name o f t h e c o p y r i g h t h o l d e r s nor t he names o f i t s 13 ∗c o n t r i b u t o r s may be used t o e n do rs e or promote p ro d u ct s d e r i v e d from 14 ∗t h i s s o f t w a r e w i t h ou t s p e c i f i c p r i o r w r i t t e n p e rm i s si o n . 15 ∗ 16 ∗THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS 17 ∗"AS IS " AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT 18 ∗LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR 19 ∗A PARTICULAR PURPOSE ARE DISCLAIMED . IN NO EVENT SHALL THE COPYRIGHT 20 ∗OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT , INCIDENTAL , 21 ∗SPECIAL , EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT 22 ∗LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES ; LOSS OF USE, 23 ∗DATA, OR PROFITS ; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY 24 ∗THEORY OF LIABILITY , WHETHER IN CONTRACT, STRICT LIABILITY , OR TORT 25 ∗(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE 26 ∗OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. 27 ∗ 28 ∗Au tho rs : E r ik H a l l n o r 29 ∗/ 30 31 /∗ ∗ 32 ∗@ f i l e 33 ∗D e c l a r a t i o n o f a memory t r a c e r e ad e r f o r a p in memory t r a c e . 34 ∗/ 35 36 #include " base / c a l l b a c k . hh " 37 #include " cpu/ t r a c e / r e a d e r / p in_r ead er . hh " 38 #include " sim / s im _exit . hh " 39 #include <s e t > 40 41 //TODO: l o o k why t h e u se r i n t e r r u p t r e c e i v e d e v en t doe sn ’ t c a l l s t he C a l l b a c k 42 43 /∗ ∗ C a l l b a c k t o c l e a n t h e qu eu es ∗/ 44 class DeleteQueuesCallback : public Callback { 45 public : 46 DeleteQueuesCallback (); 47 void process (); 48 } ; 49 s t a t i c DeleteQueuesCallback dqc ; 50 51 52 /∗ ∗ L i s t o f PinReader e le m e n t s f o r t h e queue d e l e t i n g c a l l b a c k ∗ ∗ / 53 s t a t i c s t d : : s et <PinReader ∗> readers ; 54 55 DeleteQueuesCallback :: DeleteQueuesCallback () { 56 registerExitCallback(t h i s ) ; 57 } 58 59 void DeleteQueuesCallback :: process () { 60 for ( s t d : : s et <PinReader ∗>:: i t e r a t o r i t = r e a d e r s . b egi n ( ) ; i t != r e a d e r s . end ( ) ; i t ++) { 61 (∗i t )−>removeQueues (); 62 } 63 } 64 65 //TODO: Send c l i e n t f i n a l i z a t i o n e v e n t s i f n e c e s sa r y 66 67 void PinReader : : removeQueues ( ) { 68 i f ( q ) { 69 client_fini2(q); 70 q = NULL; 71 } 72 i f ( i q ) { 73 c l i e n t _ f i n i ( iq ) ; 74 i q = NULL ; 75 } 76 warn( " Done " ) ; 77 } 78 79 PinReader : : PinReader (const PinReaderParams ∗p ) : MemTraceReader ( p ) , s i m u l a t i n g ( false) { 80 i q = c l i e n t _ i n i t ( ) ; 81 q = c l i e n t _ i n i t 2 ( ) ; 82 // Wait f o r t he i n i t i a l e v e n t 54
Appendix B. Relevant source code 83 while ( q−>empty ( ) ) { YIELD ( ) ; } 84 drop = true ; 85 r e a d e r s . i n s e r t ( t h i s ) ; 86 } 87 88 PinReader : :~ PinReader () { 89 removeQueues () ; 90 r e a d e r s . e r a s e ( t h i s ) ; 91 } 92 93 94 MemTraceRequestPtr PinReader : : getNextRequest ( MemTraceReader : : re as on &r eas on ) 95 { 96 MemCmd : : Command cmd ; 97 MemTraceRequestPtr req ; 98 Request : : F la gs f l a g s ; 99 i f ( drop ) { 100 a s s e r t ( ! q−>empty ( ) ) ; 101 q−>pop ( ) ; // Drop p r e v i o u s d at a 102 } 103 while (true ) { 104 // Wait f o r new t r a c e s i f t h e s e r v e r d ie d j u s t sen d NULL 105 while ( q−>wait_empty_cond ( ) && iq−>wait_empty_cond ( ) ) YIELD ( ) ; 106 //TODO: t h i s s t i l l n eed s some c le a ni n g , t he CPU must end any a c c e s s e s b e f o r e t he r e s e t , same b e f o r e t h e dump 107 switch ( q−>r e c e i v e _ c o n t r o l ( ) ) { 108 case SERVER_DIED: 109 // The l a s t dump s h o u l d be made by m5 i t s e l f 110 re aso n = MemTraceReader : :EOT; 111 drop = false ; 112 return NULL; 113 case SERVER_SIM_END: 114 simulating = false ; 115 q−>ack_control (); 116 re aso n = MemTraceReader : :STAT_DUMP; 117 drop = false ; 118 return NULL; 119 case SERVER_SIM_START: 120 simulating = true ; 121 q−>ack_control (); 122 re as on = MemTraceReader : : STAT_RESET; 123 drop = false ; 124 return NULL; 125 case NONE: 126 break ; 127 default : 128 warn( " S ta t e ␣ not ␣ s uppor ted ! " ) ; 129 } 130 i f ( ! q−>empty ( ) ) { 131 switch ( q−>g e t t a i l ( ) . getType ( ) ) { 132 case ACCMEM: { 133 const MemAccess &ma = q−>g e t t a i l ( ) . getCMa ( ) ; 134 switch (ma . getType ( ) ) { 135 case ACCEXEC: 136 f l a g s . s e t ( Request : : INST_FETCH ) ; 137 cmd = MemCmd : : ReadReq ; 138 break ; 139 case ACCREAD: 140 cmd = MemCmd : : ReadReq ; 141 break ; 142 case ACCWRITE: 143 cmd = MemCmd : : WriteReq ; 144 break ; 145 case ACCPREFETCH: 146 f l a g s . s e t ( Request : : PREFETCH) ; 147 cmd = MemCmd : : ReadReq ; 148 break ; 149 default : 150 pa nic ( " Access ␣ type ␣unknown " ) ; 151 } 152 Addr ea = ( Addr )ma. getEA ( ) ; 153 ea &= ( Addr ) 13 42 17 72 7; // 128Mb −1 : P 154 //By d e f a u l t ti me i s s e t t o 0 155 req = new MemTraceRequest ( ( Addr ) ea , ( int )ma . g e t S i z e ( ) , f l a g s , cmd ) ; 156 drop = true ; 157 return re q ; 158 } 159 case STARTTH: 160 case INVALDATA: 161 default : 162 pa nic ( " Unexpected ␣ data ␣ type " ) ; 163 } 164 } 165 while ( ! iq−>empty ( ) ) { 55
Appendix B. Relevant source code 249 }else { 250 //TODO: h a n dle s t a t s 251 // i f ( p kt −>is Read ( ) ) { 252 // numReads++; 253 // numReadsStat++; 254 // } e l s e { 255 // a s s e r t ( pkt −>i s W r i t e ( ) ) ; 256 // numWrites++; 257 // numWritesStat++; 258 // } 259 } 260 261 pkt−>del et e Da ta ( ) ; 262 delete pkt−>r eq ; 263 delete pkt ; 264 i f ( ! t ick Eve nt . s ched ul ed ( ) ) 265 s c h e d u l e (& t ick Eve nt , max( c ur Tick ( ) + t i c k s ( 1 ) , ( n extRequ est ? nextReq ues t−>time ( ) : 0 ) ) ) ; 266 } 267 268 TraceCPU ∗ 269 TraceCPUParams : : c r e a t e ( ) 270 { 271 return new TraceCPU( t h i s ) ; 272 } 273 274 /∗To convert∗/ 275 276 277 // v o i d 278 // MemTest : : c o m p l e t e R e q u e s t ( P a c k e t Pt r p k t ) 279 // { 280 // Re que st ∗r e q = p kt −>r e q ; 281 // 282 // i f ( issueDmas ) { 283 // dmaOutstanding = f a l s e ; 284 // } 285 // 286 // DPRINTF( MemTest , " c o m p l e t i n g %s a t a d d r e s s %x ( b l k %x ) %s \n " , 287 // pk t −>i s W r i t e ( ) ? " w r i t e " : " r ea d " , 288 // req−>ge tP ad dr ( ) , b lo c kA dd r ( req −>ge tPad dr ( ) ) , 289 // pk t −>i s E r r o r ( ) ? " e r r o r " : " s u c c e s s " ) ; 290 // 291 // MemTestSenderState ∗s t a t e = 292 // dynamic_cast<MemTestSenderState ∗>( p kt −>senderState ); 293 // 294 // u i n t 8 _ t ∗da ta = s t a t e −>da t a ; 295 // u i n t 8 _ t ∗p kt_ da ta = pk t −>g e t P t r <u in t8 _t > ( ) ; 296 // 297 // //Remove t h e a d d r e s s from t h e l i s t o f o u t s t a n d i n g 298 // s t d : : s e t <un si gn ed >: : i t e r a t o r removeAddr = 299 // o u t s t a n d i n g A d d r s . f i n d ( r eq −>g et P ad dr ( ) ) ; 300 // a s s e r t ( removeAddr != o u t s t a n d i n g A d d r s . end ( ) ) ; 301 // o u t s t a n d i n g A d d r s . e r a s e ( removeAddr ) ; 302 // 303 // i f ( p kt −>i s E r r o r ( ) ) { 304 // if (! suppress_func_warnings) { 305 // warn ( " F u n c t i o n a l A cc ess f a i l e d f o r %x a t %x \n " , 306 // pk t −>i s W r i t e ( ) ? " w r i t e " : " r ea d " , r eq−>ge tPa dd r ( ) ) ; 307 // } 308 // } e l s e { 309 // i f ( p kt −>is Read ( ) ) { 310 // i f (memcmp( pk t_data , d ata , p kt −>g e t S i z e ( ) ) != 0) { 311 // pan ic ("% s : r ea d o f %x ( b l k %x ) @ c y c l e %d " 312 // " r e t u r n s %x , e x p e c t e d %x \n " , name ( ) , 313 // req−>ge tP ad dr ( ) , b lo c kA dd r ( req −>ge tPad dr ( ) ) , cu rT ick ( ) , 314 // ∗pkt_data , ∗da t a ) ; 315 // } 316 // 317 // numReads++; 318 // numReadsStat++; 319 // 320 // i f ( numReads == ( u i n t 6 4_ t ) n e x t P r o g r e s s M e s s a g e ) { 321 // c c p r i n t f ( c err , "%s : c om pl et ed %d read , %d w r i t e a c c e s s e s @%d\n " , 322 // name ( ) , numReads , numWrites , c ur T ic k ( ) ) ; 323 // n e x t P r o g r e s s M e s s a g e += p r o g r e s s I n t e r v a l ; 324 // } 325 // 326 // i f ( maxLoads != 0 && numReads >= maxLoads ) 327 // ex itS im Loo p ( " maximum number o f l o a d s r ea ch ed " ) ; 328 // } e l s e { 329 // a s s e r t ( pkt −>i s W r i t e ( ) ) ; 330 // f u nc Po r t . w r i t e B l o b ( req −>g etP ad dr ( ) , p kt_dat a , req −>g e t S i z e ( ) ) ; 331 // numWrites++; 62
Appendix B. Relevant source code 332 // numWritesStat++; 333 // } 334 // } 335 // 336 // n oR es po ns eCy c les = 0 ; 337 // d e l e t e s t a t e ; 338 // d e l e t e [ ] d a ta ; 339 // d e l e t e pkt −>r e q ; 340 // d e l e t e p k t ; 341 // i f ( ! t i c k E v e n t . s c h e d u l e d ( ) ) 342 // s c h e d u l e ( t i c k E v e n t , c u rT ic k ( ) + t i c k s ( 1 ) ) ; 343 // } 344 345 // v o i d 346 // MemTest : : t i c k ( ) 347 // { 348 // 349 // //make new r e q u e s t 350 // /∗u n s i g n e d cmd = random ( ) % 1 0 0 ; 351 // ∗u n s i gne d o f f s e t = random ( ) % s i z e ; 352 // ∗u n s i g n e d b a s e = random ( ) % 2 ; 353 // ∗u i n t6 4 _ t d a ta = random ( ) ; 354 // ∗u n si gn ed a c c e s s _ s i z e = random ( ) % 4 ; 355 // ∗b o o l u n c a c h e a b l e = ( random ( ) % 100 ) < p e r c e n t U n c a c h e a b l e ; 356 // ∗ 357 // ∗u n s i g n e d dm a_ac cess _si ze = random ( ) % 4 ; ∗/ 358 // u n s i g n e d cmd = 0 ; 359 // o f f s e t ++; 360 // u n s i g n e d b a s e = 0 ; 361 // u int 6 4_t d ata = random ( ) ; 362 // u ns ign e d a c c e s s _ s i z e = 0 ; 363 // b o o l u n c a c h e a b l e = f a l s e ; 364 // 365 // u n s i g n e d d ma_ acc ess_ size = random ( ) % 4 ; 366 // 367 // // I f we aren ’ t doi n g c o p i e s , us e i d as o f f s e t , and do a f a l s e s h ar i n g 368 // //mem t e s t e r 369 // //We can e l i m i n a t e t h e l o w e r b i t s o f t h e o f f s e t , and t h en u se t h e i d 370 // // t o o f f s e t w i t h i n t h e b l k s 371 // // o f f s e t = b l ock A ddr ( o f f s e t ) ; 372 // // o f f s e t += i d ; 373 // // a c c e s s _ s i z e = 0 ; 374 // // dma_access_size = 0; 375 // 376 // Re que st ∗r e q = new R eq u es t ( ) ; 377 // R eq ue s t : : F l a g s f l a g s ; 378 // Addr pa ddr ; 379 // 380 // i f ( u n c a c h e a b l e ) { 381 // f l a g s . s e t ( Req ue st : : UNCACHEABLE) ; 382 // pa ddr = uncacheAddr + o f f s e t ; 383 // } e l s e { 384 // pa ddr = ( ( bas e ) ? b aseAd dr1 : ba seAddr2 ) + o f f s e t ; 385 // } 386 // b o o l d o _ f u n c t i o n a l = f a l s e ; 387 // 388 // i f ( issueDmas ) { 389 // p addr &= ~ ( ( 1 << dma _acce ss_si ze ) −1 ) ; 390 // req−>s e t P h y s ( padd r , 1 << d ma_access_size , f l a g s ) ; 391 // req−>s e t T h r e a d C o n t e x t ( i d , 0 ) ; 392 // } e l s e { 393 // pa ddr &= ~ ( ( 1 << a c c e s s _ s i z e ) −1 ) ; 394 // req−>s e t P h y s ( paddr , 1 << a c c e s s _ s i z e , f l a g s ) ; 395 // req−>s e t T h r e a d C o n t e x t ( i d , 0 ) ; 396 // } 397 // a s s e r t ( req −>g e t S i z e ( ) == 1 ) ; 398 // 399 // uint8_t ∗r e s u l t = new u i n t 8 _ t [ 8 ] ; 400 // 401 // i f (cmd < p e rc e n tR e ads ) { 402 // // rea d 403 // 404 // // For now we o nly a l l o w one o u t s t a n d i n g r e q u e s t p er a d d re s s 405 // // p er t e s t e r T his means we assume CPU d oe s w r i t e f o r w a r d i n g 406 // // t o r e a d s t h a t a l i a s s om e th i ng i n t h e cpu s t o r e b u f f e r . 407 // i f ( o u t s t a n d i n g A d d r s . f i n d ( p addr ) != o u t s t a n d i n g A d d r s . end ( ) ) { 408 // d e l e t e [ ] r e s u l t ; 409 // d e l e t e r eq ; 410 // return ; 411 // } 412 // 413 // o u t s t a n d i n g A d d r s . i n s e r t ( pa ddr ) ; 414 // 63
Appendix B. Relevant source code 415 // // ∗∗∗∗∗ NOTE FOR RON: I ’m n o t s u r e how t o a c c e s s checkMem . −Kevin 416 // f u n c P o r t . r e a d B lo b ( r eq −>ge tPa dd r ( ) , r e s u l t , req −>g e t S i z e ( ) ) ; 417 // 418 // c c p r i n t f ( ce rr , 419 // " i d %d i n i t i a t i n g %s r e a d a t addr %x ( b l k %x ) e x p e c t i n g %x\n " , 420 // id , d o _ f u n c t i o n a l ? " f u n c t i o n a l " : " " , req −>g etP ad dr ( ) , 421 // b l ock A ddr ( req−>getPaddr ()) , ∗r e s u l t ) ; 422 // 423 // P ac k e tP t r p k t = new P ac ke t ( re q , MemCmd : : ReadReq , P ac ke t : : B ro a dc as t ) ; 424 // p kt −>s e t S r c ( 0 ) ; 425 // p kt −>d at aDy nam icAr ray ( new u i n t 8 _ t [ re q−>g e t S i z e ( ) ] ) ; 426 // MemTestSenderState ∗s t a t e = new MemTestSenderState ( r e s u l t ) ; 427 // p kt −>s e n d e r S t a t e = s t a t e ; 428 // 429 // i f ( d o _ f u n c t i o n a l ) { 430 // a s s e r t ( pkt −>n ee ds Re sp onse ( ) ) ; 431 // pk t −>s e t S u p p r e s s F u n c E r r o r ( ) ; 432 // c a ch e Po r t . s e n d F u n c t i o n a l ( p k t ) ; 433 // completeRequest(pkt ); 434 // } e l s e { 435 // s en dP kt ( p k t ) ; 436 // } 437 // } e l s e { 438 // // w r i t e 439 // 440 // // For now we o n l y a l l o w one o u t s t a n d i n g r e q u e s t p er a d d r e e s s 441 // // p er t e s t e r . Thi s means we assume CPU d oe s w r i t e f o r w a r d i n g 442 // // t o r e a d s t h a t a l i a s s om e th i ng i n t h e cpu s t o r e b u f f e r . 443 // i f ( o u t s t a n d i n g A d d r s . f i n d ( p addr ) != o u t s t a n d i n g A d d r s . end ( ) ) { 444 // d e l e t e [ ] r e s u l t ; 445 // d e l e t e r eq ; 446 // return ; 447 // } 448 // 449 // o u t s t a n d i n g A d d r s . i n s e r t ( pa ddr ) ; 450 // 451 // DPRINTF(MemTest , " i n i t i a t i n g %s w r i t e a t addr %x ( b l k %x ) v a l u e %x\n " , 452 // d o _ f u n c t i o n a l ? " f u n c t i o n a l " : " " , r eq−>getPaddr () , 453 // blo c k Ad d r ( req −>g etPa dd r ( ) ) , d at a & 0 x f f ) ; 454 // 455 // P ac k e tP t r p k t = new P ac ke t ( re q , MemCmd : : WriteReq , P ac ke t : : B ro ad ca s t ) ; 456 // p kt −>s e t S r c ( 0 ) ; 457 // uint8_t ∗p k t_ d at a = new u i n t 8 _ t [ req −>g e t S i z e ( ) ] ; 458 // p kt −>dataDynamicArray ( p kt _d at a ) ; 459 // memcpy ( p kt_ data , &d ata , r eq −>g e t S i z e ( ) ) ; 460 // MemTestSenderState ∗s t a t e = new MemTestSenderState ( r e s u l t ) ; 461 // p kt −>s e n d e r S t a t e = s t a t e ; 462 // 463 // i f ( d o _ f u n c t i o n a l ) { 464 // pk t −>s e t S u p p r e s s F u n c E r r o r ( ) ; 465 // c a ch e Po r t . s e n d F u n c t i o n a l ( p k t ) ; 466 // completeRequest(pkt ); 467 // } e l s e { 468 // s en dP kt ( p k t ) ; 469 // } 470 // } 471 // } 64
Appendix B. Relevant source code pintrace.py 1# C o p y r i g h t ( c ) 2006 −2007 The Regents o f The U n i v e r s i t y o f M ichigan 2# A l l r i g h t s r e s e r v e d . 3# 4# R e d i s t r i b u t i o n and use in s ou rc e and b i na r y forms , w it h or w it h ou t 5# m o d i f i c a t i o n , a re p e r m i t t e d p r o v i d e d t h a t t h e f o l l o w i n g c o n d i t i o n s a re 6# met : r e d i s t r i b u t i o n s o f s o u r c e co de must r e t a i n t h e ab ov e c o p y r i g h t 7# n o t ic e , t h i s l i s t o f c o n d i t i o n s and t he f o l l o w i n g d i s c l a i m e r ; 8# r e d i s t r i b u t i o n s i n b i n a ry form must r e p r odu c e t h e ab ov e c o p y r i g h t 9# n o t ic e , t h i s l i s t o f c o n d i t i o n s and t he f o l l o w i n g d i s c l a i m e r i n t he 10 # d oc ume nt at io n and / or o t h er m a t e r i a l s p r ov id ed w i t h t he d i s t r i b u t i o n ; 11 # n e i t h e r t he name o f t h e c o p y r i g h t h o l d e r s nor t he names o f i t s 12 # c o n t r i b u t o r s may be us ed t o e nd or s e or promote p r od u c t s d e r i v e d from 13 # t h i s s o f t w a r e w i t ho u t s p e c i f i c p r i o r w r i t t e n p ermi ss i on . 14 # 15 # THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS 16 # "AS IS " AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT 17 # LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR 18 # A PARTICULAR PURPOSE ARE DISCLAIMED . IN NO EVENT SHALL THE COPYRIGHT 19 # OWNER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT , INCIDENTAL , 20 # SPECIAL , EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT 21 # LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES ; LOSS OF USE, 22 # DATA, OR PROFITS ; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY 23 # THEORY OF LIABILITY , WHETHER IN CONTRACT, STRICT LIABILITY , OR TORT 24 # (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE 25 # OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. 26 # 27 # Au tho rs : Ron D r e s l i n s k i 28 29 import optparse 30 import sys 31 32 import m5 33 from m5. objects import ∗ 34 35 p a r s e r = o pt p ar s e . Opt ionPa rser ( ) 36 37 #p a r s e r . a dd_o ptio n ("−m" , "−− m axt i ck " , t yp e =" i n t " , d e f a u l t =m5. MaxTick , 38 #m etav ar ="T" , 39 #h e l p =" St op a f t e r T t i c k s " ) 40 41 ( o p ti on s , a r g s ) = p a r s e r . pa rse _ar gs ( ) 42 43 i f a r g s : 44 print " E rror : ␣ s c r i p t ␣ doesn ’ t ␣ tak e ␣any␣ p o s i t i o n a l ␣ arguments " 45 sys . e x i t ( 1 ) 46 47 # d e f i n e p r o t o t y p e L1 ca che 48 proto_l 1 = BaseCache ( s i z e = ’ 32kB ’ , a s s o c = 4 , b l o c k _ s i z e = 128 , 49 l a t e n c y = ’ 1 ns ’ , tgts_per_mshr = 1) 50 51 proto_l 1 . mshrs = 1 52 53 pr = PinReader ( ) 54 55 tcpu = TraceCPU ( t r a c e = pr ) 56 57 # n e x t comes L1 c ache , i f any 58 #p r o t o t y p e s . i n s e r t ( 0 , p r oto _l1 ) 59 60 # s ys tem s i m u l a t e d 61 62 system = System (physmem = PhysicalMemory ( l a t e n c y = " 100 ns " ) ) 63 64 new_bus = Bus ( c l o c k=" 500MHz" , width =16) 65 system . physmem . cpu_side_bus = new_bus 66 system . physmem . port = new_bus . master 67 68 data_l1 = BaseCache ( s i z e = ’ 32kB ’ , a s s o c = 4 , b l o c k _ s i z e = 64 , 69 l a t e n c y = ’ 1 ns ’ , tgts_per_mshr = 8) 70 data_l1 . mshrs = 1 71 72 ins _ l1 = BaseCache ( s i z e = ’ 32kB ’ , a s s o c = 4 , b l o c k _ s i z e = 64 , 73 l a t e n c y = ’ 1 ns ’ , tgts_per_mshr = 8) 74 ins _ l1 . mshrs = 1 75 76 new_bus . cach e = [ data_l1 , i n s_ l 1 ] 77 new_bus . s l a v e = data_l1 . mem_side 78 new_bus . s l a v e = i ns _ l1 . mem_side 79 80 data_l1 . cpu = tcpu 81 82 tcpu . data = data_l1 . cpu_side 65
Appendix B. Relevant source code 83 tcpu . i n s t r u c t i o n s = in s_l 1 . cpu_side 84 #d e f m a ke _l e ve l ( s pec , p r o t o t y p e s , a t ta c h _o b j , a t t a c h _ p o r t ) : 85 #f a n o u t = s p e c [ 0 ] 86 #p a r e nt = a t t a c h _ o b j # u se a t t a c h o b j as c o n f i g p a r e nt t o o 87 # i f l e n ( s p e c ) > 1 and ( f a n o u t > 1 or o p t i o n s . f o r c e _ b u s ) : 88 #new_bus = Bus ( c l o c k ="500MHz" , w id th =16) 89 #new_bus . p o r t = g e t a t t r ( a t t a c h _ o b j , a t t a c h _ p o r t ) 90 #p a r e n t . cpu_ side_ bus = new_bus 91 #a t t a c h _ o b j = new_bus 92 #a t t a c h _ p o r t = " p o r t " 93 #o b j s = [ p r o t o t y p e s [ 0 ] ( ) f o r i i n x ra nge ( fa n o ut ) ] 94 # i f l e n ( s p e c ) > 1 : 95 ## we j u s t b u i l t c ac hes , more l e v e l s t o go 96 #p a r e n t . c ac he = o b j s 97 #f o r c ach e i n o b j s : 98 #c a c h e . mem_side = g e t a t t r ( a t t a c h _ o b j , a t t a c h _ p o r t ) 99 #ma ke _l evel ( s pe c [ 1 : ] , p r o t o t y p e s [ 1 : ] , cache , " c pu_si de ") 100 #e l s e : 101 ## we j u s t b u i l t t h e MemTest o b j e c t s 102 #p a r e n t . cpu = o b j s 103 #f o r t in o b j s : 104 #t . t e s t = g e t a t t r ( a t t a c h _ o b j , a t t a c h _ p o r t ) 105 #t . f u n c t i o n a l = sys te m . funcmem . p o rt 106 107 #m ak e_ le ve l ( t r e e s p e c , p r o t o t y p e s , s yst em . physmem , " p o r t " ) 108 109 #−−−−−−−−−−−−−−−−−−−−−−− 110 # run simulation 111 #−−−−−−−−−−−−−−−−−−−−−−− 112 113 r o o t = Root ( f u l l _ s y s t e m = Fa lse , syst em = s ys te m ) 114 r o o t . syst em . mem_mode = ’ timi ng ’ 115 116 r o o t . syst em . sy st em_por t = r o o t . syst em . physmem . p o r t 117 118 # Not much p o i n t i n t h i s b e i n g h i g h e r th an t h e L1 l a t e n c y 119 m5. t i c k s . se tG lo ba lF requency ( ’ 1 ns ’ ) 120 121 # i n s t a n t i a t e c o n f i g u r a t i o n 122 m5. i n s t a n t i a t e ( ) 123 124 # s i m u l a t e u n t i l program t e r m i n a t e s 125 e x it _e v en t = m5 . s i m u l a t e (m5 . MaxTick ) 126 127 print ’ E x i t i n g ␣@␣ t i c k ’ , m5. cur Tick ( ) , ’ be cau se ’ , e xi t_e ve nt . getCause ( ) 66
Bibliography [1] Kenneth Barr. Dinerotool. Oct. 2005. url:http://kbarr.net. [2] Nathan Binkert et al. “The gem5 simulator”. In: SIGARCH Comput. Archit. News 39.2 (Aug. 2011), pp. 1–7. issn: 0163-5964. doi:10.1145/ 2024716.2024718.url:http://doi.acm.org/10.1145/2024716. 2024718. [3] Zhongliang Chen et al. The Multi2Sim Simulation Framework.url: http://www.multi2sim.org/files/multi2sim-r277.pdf. [4] Circular buffer. Nov. 2012. url:http : / / en . wikipedia . org / w / index.php?title=Circular_buffer&oldid=522370238#Always_ Keep_One_Slot_Open. [5] H. J. Curnow and B. A. Wichmann. “A synthetic benchmark”. In: The Computer Journal 19.1 (1976), pp. 43–49. doi:10.1093/comjnl/19. 1.43. eprint: http://comjnl.oxfordjournals.org/content/19/ 1/43.full.pdf+html.url:http://comjnl.oxfordjournals.org/ content/19/1/43.abstract. [6] Susan L. Graham, Peter B. Kessler, and Marshall K. Mckusick. “Gprof: A call graph execution profiler”. In: SIGPLAN Not. 17.6 (June 1982), pp. 120–126. issn: 0362-1340. doi:10.1145/872726.806987.url: http://doi.acm.org/10.1145/872726.806987. 67
Bibliography Bibliography [7] Mark Hill and Jan Edler. Dinero IV Trace-Driven Uniprocessor Cache Simulator. Feb. 1998. url:http://www.cs.wisc.edu/~markhill/ DineroIV/. [8] Chi-Keung Luk et al. “Pin: building customized program analysis tools with dynamic instrumentation”. In: Proceedings of the 2005 ACM SIGPLAN conference on Programming language design and implementation. PLDI ’05. Chicago, IL, USA: ACM, 2005, pp. 190–200. isbn: 1-59593-056-6. doi:10.1145/1065010.1065034.url:http://doi. acm.org/10.1145/1065010.1065034. [9] J.E. Miller et al. “Graphite: A distributed parallel simulator for multicores”. In: High Performance Computer Architecture (HPCA), 2010 IEEE 16th International Symposium on. Jan. 2010, pp. 1 –12. doi: 10.1109/HPCA.2010.5416635.url:http://groups.csail.mit. edu/carbon/docs/graphite_hpca2010_preprint.pdf. [10] Vijay Janapa Reddi et al. “PIN: a binary instrumentation tool for computer architecture research and education”. In: Proceedings of the 2004 workshop on Computer architecture education: held in conjunction with the 31st International Symposium on Computer Architecture. WCAE ’04. Munich, Germany: ACM, 2004. doi:10.1145/1275571.1275600. url:http://doi.acm.org/10.1145/1275571.1275600. [11] Cloyce D. Spradling. “SPEC CPU2006 Benchmark Tools”. In: SIGARCH Computer Architecture News 35 (1 Mar. 2007). [12] Richard M. Stallman. GDB manual: the GNU source-level debugger. 2nd, GDB version 2.5. Free Software Foundation, Inc. 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301, USA, Tel: (617) 876-3296, Feb. 1988, pp. ii + 63. 68
Bibliography Bibliography [13] Richard M. Stallman. Using and Porting GNU CC. Tech. rep. 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301, USA, Tel: (617) 876-3296: Free Software Foundation, Inc., 1988. [14] The gcc website.url:http://gcc.gnu.org/. [15] The gdb website.url:http://www.gnu.org/software/gdb/. [16] The Gem5 website.url:http://www.gem5.org/. [17] The gprof website.url:http://sourceware.org/binutils/docs/ gprof/. [18] The Graphite website.url:http://groups.csail.mit.edu/carbon/ ?page_id=111. [19] The modified SPLASH-2 website.url:www.capsl.udel.edu/splash/. [20] The Multi2Sim website.url:http://www.multi2sim.org/. [21] The Pin website.url:http://software.intel.com/en-us/articles/ pintool/. [22] The SPEC CPU2006 website.url:http://www.spec.org/cpu2006/. [23] The SPLASH-2 website.url:http://web.archive.org/web/http: //www-flash.stanford.edu/apps/SPLASH/. [24] vanDooren. Creating a thread safe producer consumer queue in C++ without using locks. Jan. 2007. url:http: / / msmvps . com / blogs / vandooren / archive / 2007 / 01 / 05 / creating - a - thread - safe - producer-consumer-queue-in-c-without-using-locks.aspx. [25] Reinhold P. Weicker. “Dhrystone: a synthetic systems programming benchmark”. In: Commun. ACM 27.10 (Oct. 1984), pp. 1013–1030. issn: 0001-0782. doi:10.1145/358274.358283.url:http://doi. acm.org/10.1145/358274.358283. 69
Bibliography Bibliography [26] S.C. Woo et al. “The SPLASH-2 Programs: Characterization and Methodological Considerations”. In: Proc. of the 22nd International Symposium on Computer Architecture. June 1995. 70