scieee AI-readable full text Open interactive document viewer

On the Precision of Dynamic Program Fingerprints based on Performance Counters

Faustino da Silva, Anderson

Abstract

Task classification is the challenge of determining whether two binary programs perform the same task. This problem is essential in scenarios such as malware identification, plagiarism detection, and redundancy elimination. Classification can be performed statically or dynamically. In the former case, the classifier analyzes the binary image of the program, whereas in the latter it observes the program's execution. Recent research has demonstrated that dynamic classification is more accurate, particularly in adversarial settings where programs may be obfuscated. This remains true even when both classifiers use the exact representation of programs, such as histograms of instruction opcodes. The superior accuracy of dynamic classification stems from its ability to disregard dead code inserted during the obfuscation process.However, state-of-the-art dynamic techniques, such as Valgrind plugins, can slow down program execution by as much as 100 times due to binary instrumentation.This is the artifact of our paper that proposes to eliminate this overhead by replacing program instrumentation with the sampling of hardware performance counters.

Full text

On the Precision of Dynamic Program Fingerprints based on Performance Counters Anderson Faustino da Silva UEM Maring´ a, Brazil afsilv[email protected] Marcelo Borges Nogueira UFRN Natal, Brazil [email protected] S´ ergio Queiroz de Medeiros UFRN Natal, Brazil ser[email protected] Jeronimo Castrillon TU Dresden Dresden, Germany [email protected] Fernando Magno Quint˜ ao Pereira UFMG Belo Horizonte, Brazil [email protected] Abstract—Task classification is the challenge of determining whether two binary programs perform the same task. This problem is important in scenarios such as malware identification, plagiarism detection, and redundancy elimination. Classification can be performed statically or dynamically. In the former case, the classifier analyzes the binary image of the program, whereas in the latter it observes the program’s execution. Recent research has demonstrated that dynamic classification is more accurate, particularly in adversarial settings where programs may be obfuscated. This remains true even when both classifiers use the same representation of programs, such as histograms of instruction opcodes. The superior accuracy of dynamic classification stems from its ability to disregard dead code inserted during obfuscation. However, state-of-the-art dynamic techniques, such as Valgrind plugins, might slow down program execution by as much as 100x due to binary instrumentation. This paper proposes to eliminate this overhead entirely by replacing program instrumentation with the sampling of hardware performance counters. Our findings show advantages and limitations in this approach. On the positive side, classifiers based on hardware counters impose almost no runtime overhead while retaining greater accuracy than purely static classifiers, particularly in the presence of obfuscation. On the downside, counter-based classifiers are slightly less accurate than instrumentation-based approaches and offer coarser granularity, being limited to whole-program classification rather than individual functions. Despite these limitations, our results challenge the conventional belief that dynamic code classifiers are too costly to be deployed in environments such as online servers, operating systems, and virtual machines. Index Terms—Security, Binary Diffing, Code Classification DATA AVAILABILITY STATEMENT The reproducible artifact for this paper is available at either https://github.com/ComputerSystemsLaboratory/Rouxinol or via Zenodo [1]. ACKNOWLEDGMENT This work was partially funded by the AI competence center ScaDS.AI Dresden/Leipzig in Germany (01IS18026A-D) and the German Research Council (DFG) through the TransT project (552689849). ARTIFACT A. Abstract This artifact reproduces the experiments conducted in Section IV. A package containing scripts to generate the paper results and automatically produce Figures 5, 6, 7, 8, and 9 is available at https://doi.org/10.5281/zenodo.17802066 . Additionally, the artifact contains preprocessed data. In this case, it is possible to skip data generation. B. Artifact check-list • Program: VALGRIND with CFGGRIND , OBFUSCATOR-LLVM, PERF, ANACONDA, and ROUXINOL. •Compilation: o-llvm. •Dataset: CodeNet [2]. •Run-time environment: Linux with an x86-64 architecture. •Hardware: Any hardware that supports perf. •Metrics: Accuracy and runtime overhead. • Output: LLVM IR, histograms, binaries, and Figures 5, 6, 7, 8, and 9. • Experiments: Generate LLVM IR, histograms, binaries, and the figures. • How much disk space required: The experiment requires ∼50Gb. • How much time is needed to prepare workflow: To build the infraestructure requires ∼4 hours. • How much time is needed to complete experiments: The experiment, per machine, takes ∼30 days. • Publicly available: It is publicly available at https://github.com/ComputerSystemsLaboratory/Rouxinol and the artifact is available at https://doi.org/10.5281/zenodo.17802066. •Workflow framework used: Anaconda. C. Description 1) How delivered: Zenodo ( https://doi.org/10.5281/zenodo.17802066 ) and GitHub (https://github.com/ComputerSystemsLaboratory/Rouxinol). D. Installation 1) Download the artifacts from Zenodo (https://doi.org/10.5281/zenodo.17802066) and decompress it: $ tar xfJ cgo2026.tar.xz 2) Download and install Anaconda. 3) Create rouxinol environment. $ cd CGO2026/Rouxinol/pkgs $ conda create env -f conda_[cpu|gpu].yml 4) Build Rouxinol. $ cd CGO2026/Rouxinol $ ./setup.py build 5) Install VALGRIND with CFGGRIND and OBFUSCATOR-LLVM. $ cd CGO2026/Rouxinol $ ./install_deps.sh 6) Install perf. E. Experiment workflow 1) Activate rouxinol environment. $ conda activate rouxinol 2) Edit run.sh. You need to set TOP DIR to point to the CGO2026 directory. In addition, set the other variables accordingly. $ cd CGO2026/artifact $ vim run.sh 3) Run the experiments. • Run the complete experiment (generate the data, run classification, and plot the figures). $ cd CGO2026/artifact $ ./run.sh • Skip data generation (run classification and plot the figures). $ cd CGO2026/artifact $ ./run_skip_generator.sh •Only generate the figures. $ cd CGO2026/artifact $ ./run_only_figures.sh F. Details •Anaconda : The bash scripts consider that the Anaconda directory is located inside CGO2026. If it is not true, edit CONDA DIRECTORY in run.sh, run skip generator.sh, and run only figures.sh. •Rouxinol directory : The variable ROUXINOL DIRECTORY points to the ROUXINOL directory. •Rouxinol data directory : The variable ROUXINOL DATA DIRECTORY points to the directory that contains the dataset (32 problems, each with 100 solutions). This directory is CGO2026/rouxinol data. •Output directory : The variable ROUXINOL OUTPUT DIRECTORY points to the directory where data and figures will be generated. In run.sh it points to TOP DIR/output/A1, indicating that the script will generate data for the machine A1. Change A1 to the label of the other machines. •Test : The experiments take forever. However, it is possible to run a small experiment (2 problems each with 10 solutions): $ cd CGO2026/artifact $ ./run.sh test •Data generation : It is possible to skip data generation (LLVM IR, histogram, and binary). In this case, we provide our preprocessed data (CGO20226/preprocessed). Edit the variable ROUXINOL OUTPUT DIRECTORY to point to the preprocessed directory and the specific machine (A1, A2, or I1). The script run skip generator.sh will generate the statistics for the existing directories and match O1 that already exist. Therefore, be careful not to lose preprocessed data. G. Evaluation and expected result After completion, everything produced can be found in the following directories: •bcf|fla|O0|O1|O2|O3|ollvm|sub : LLVM IR, binaries and histograms. •statistics : classification results (yml files) and Figures 5, 7, and 8 (figure[X].pdf). •match_O1 : match data and Figure 9 (figure9 [representation].pdf). •CGO2026: Figure 6 (figure6.pdf). It is essential to note that we are considering the directories to be located within the ROUXINOL OUTPUT DIRECTORY, except CGO2026. REFERENCES [1] Anderson Faustino da Silva. On the precision of dynamic program fingerprints based on performance counters. https://doi.org/10. 5281/zenodo.17802066, November 2025. [Online; accessed 03Dec-2025]. [2] project codenet, ufinkler, Geert Janssen, VlZolotov, Fred Reiss, Jie Chen, Mihir Choudhury, Steve Martinelli, Veronika Thost, giacomo domeniconi, lindseydecker, lucaburatti7, and Ruchir Puri. Ibm/project codenet: Initial release 1.0 (may 5, 2021), May 2021.