Malware Identification by Statistical Opcode Analysis

CSEF · 2010 Mathematics & Software Honorable_mention Award

Overview

Objectives/Goals This project determined the efficacy of statistical analysis of program assembly instruction (opcode) frequencies to identify Malware from Goodware. Methods/Materials Malware and Goodware binaries were obtained and a python script was created to extract opcode frequencies from specific parts of these files. Naive Bayes models and Kmeans based models were then trained using these executables. These models were tested using a different set of programs to determine their efficacy at identifying Malware from Goodware. Results The best Naive Bayes model had a recall of 1 for Malware and .8 for Goodware. Conclusions/Discussion Differences in opcode frequencies can differentiate Malware from Goodware. Certain instructions occur much more frequently in one group than in the other; these differences can be used to identify the two types of programs.

Summary statement

This project examines models that differentiate Malware from Goodware using the frequencies of program assembly instructions.

Help received

Communicated with mentor Joshua Kroll ; Pamela Durkee proofread papers and guidance

Awards (1)

  • Honorable Mention

Competition history

  • CSEF 2010 Mathematics & Software · Entry S1602

Resources

Related projects

Closest projects by meaning, across every fair and year in the corpus.

Browse more like this

Source: California Science & Engineering Fair public projects

Save projects to your library

Sign in with Google to keep track of projects you find interesting, organized into folders. An account also raises your daily allowance for “Has this been done?”, and lets you create a key for the MCP server with a much higher limit than anonymous use. Browsing stays public.

Continue with Google