Back to Glossary
By 
October 1, 2026

What Is Fuzz Testing?

Fuzz testing is an automated software testing technique that supplies a program with large volumes of malformed, unexpected, or randomly generated input in order to find the inputs that make it fail. Failure means something observable, such as a crash, a hang, a failed assertion, a memory error, or runaway resource consumption.

Fuzzing is the same practice under a shorter name, and the two terms are used interchangeably. Neither describes fuzzy logic or fuzzy matching, which concern degrees of truth and approximate string comparison. Those fields share an English root with fuzzing and nothing more.

Where the Term Came From

Fuzz testing began as a graduate class project, when in 1988 Barton Miller set his Advanced Operating Systems students at the University of Wisconsin-Madison the task of building a random character generator and pointing it at UNIX command-line utilities to see which ones broke. The assignment grew out of an observed nuisance, that programs crashed when input arrived over a noisy dial-up line, and the tool the students wrote took its name from that noise.

The results covered 88 command-line applications across six versions of UNIX. Published in 1990, they showed between 25% and 33% of those utilities crashing or hanging. The paper identified accessing memory outside the bounds of a buffer as the single largest failure category, and suggested the method might help find security holes.

Miller's group repeated the work in 1995, when it crashed up to 40% of command-line utilities in the worst case and found the GNU and Linux versions noticeably more reliable than their commercial UNIX counterparts. The study later extended to Windows NT in 2000, Mac OS X in 2006, and Linux, FreeBSD and macOS in 2020. That last round found 9 of 74 utilities failing on Linux, 15 of 78 on FreeBSD, and 12 of 76 on macOS, with 24 distinct utilities failing across the three platforms. The researchers noted that these failure rates ran somewhat higher than in their previous studies, and that the few utilities written in Rust they tested proved no more reliable than the rest. Since the pass criterion is only whether a program crashes or hangs, a Rust program that panics on malformed input fails it as surely as one that corrupts memory.

How a Fuzzer Operates

A fuzzer needs an entry point into the program, called a harness, which is a small piece of code that accepts a byte array and passes it to the function under test. It also needs a starting corpus of valid inputs, since generating a well-formed PNG or TLS handshake by chance is not feasible.

From there the loop runs mechanically, as the fuzzer mutates an input, runs the program against it, watches for a failure signal, and records anything that produces one. Crashing inputs are then minimized to the smallest version that still reproduces the failure, which is what a developer can act on.

Mutation, Generation, and Coverage Guidance

Mutation-based fuzzing takes valid samples and corrupts them, flipping bits, truncating fields, and splicing files together. It needs almost no knowledge of the input format and wastes most attempts on inputs the parser rejects immediately.

Where a grammar exists for the format or protocol, generation-based fuzzing builds inputs from that model, producing structurally valid data carrying deliberately hostile values. It reaches deeper code, at the price of someone writing the grammar first.

Coverage-guided fuzzing changed the field by instrumenting the target at compile time, letting the fuzzer observe which branches each input reaches and keep the ones that reach new code for further mutation. Input structure is discovered by feedback rather than supplied in advance. That vantage point, watching execution from inside the process, is the same one interactive application security testing uses to trace untrusted data through a running application.

Fuzz Testing Tools

The coverage-guided engines in common use are libFuzzer, AFL++, honggfuzz, and Centipede. Language-specific harnesses extend the approach beyond C and C++, among them Jazzer for the JVM, Atheris for Python, cargo-fuzz for Rust, and a fuzzing mode built into the Go toolchain.

Sanitizers matter as much as the engine. Address, memory, and undefined-behavior sanitizers instrument the build so that out-of-bounds reads and use-after-free conditions abort immediately rather than corrupting memory silently and passing the test. A fuzzer without a sanitizer finds only the failures loud enough to crash on their own.

OSS-Fuzz runs these tools as a continuous service for open-source projects. Google launched it in 2016 after Heartbleed, a memory buffer flaw in OpenSSL that fuzzing could have caught. Its maintainers report more than 50,000 bugs found and fixed as of May 2025, among them over 13,000 vulnerabilities, spread across a thousand projects.

What Fuzz Testing Misses

Every finding depends on having an oracle, meaning a way for the fuzzer to recognize that something went wrong. Miller's criterion was deliberately crude. A program passes unless it crashes or hangs, so it may respond nonsensically, or exit quietly, and still be recorded as a pass. Crashes, hangs, and sanitizer reports provide one, and nothing else does.

A function that returns the wrong account balance exits cleanly. Broken access control, business logic errors, and information disclosure through legitimate responses all look like successful runs. Fuzzing is strong against memory safety and parser robustness, and largely blind to correctness.

Coverage plateaus as well, with a campaign reaching the code its harness and corpus can lead it to and then spending indefinite compute re-exploring the same paths. Extending it requires human work, whether a new harness, a richer corpus, or a grammar for the format.

Fuzz testing exercises inputs someone thought to generate, inside an environment built for the purpose. Production supplies inputs nobody designed, aimed at code paths no harness covers. Raven Runtime Prevention works on that side of the line, judging code by what it does as it executes and halting malicious behavior as it happens, including exploitation of defects no test campaign ever reached. See it at Runtime Prevention.

Share this post