Technology

What Is a Checksum or Hash, and Why Does SHA-256 Exist?

Or: how a computer can turn an entire file into one very peculiar fingerprint. Sixty-four hexadecimal characters can tell you an awful lot about a 100-gigabyte file.

Or: How a Computer Can Turn an Entire File Into One Very Peculiar Fingerprint

Suppose you download a 6 GB operating-system image from the internet. The download completes successfully, the file appears in your Downloads folder, and its size looks about right. But how do you know the file you received is exactly the file the developer published?

Maybe one tiny portion became corrupted during storage or transfer. Perhaps a malfunctioning drive altered something. More concerning, maybe someone deliberately replaced the legitimate file with a modified copy containing malware. You could compare your file against the original byte by byte, but that requires possessing the original — which rather defeats the purpose of downloading it.

Instead, the publisher might give you something resembling this: SHA-256: 9f86d081884c7d659a2feaa0c55ad015... That long collection of hexadecimal characters may look like technological gibberish, but it provides a remarkably powerful way to answer a simple question: did I get the same data? Welcome to the world of checksums and cryptographic hashes.

Start With the Simplest Idea

At its most basic, a checksum is a value calculated from a collection of data. You take the data, run it through an algorithm, and receive a smaller value representing it. Later, you perform the calculation again — if the result has changed, something about the underlying data probably changed too.

Imagine an absurdly simple checksum that represents letters with numbers and adds them together. CAT might produce one total while CAR produces another — a crude but useful way to notice some changes. The problem is that a checksum this simple would be terrible, because it can miss errors entirely. The sequence 10 + 20 + 30 produces 60. So does 30 + 20 + 10. The data changed; the checksum didn't. That's called a collision — two different inputs producing the same output — and every checksum or fixed-length hash must theoretically have collisions, since there are infinitely more possible inputs than outputs. The real question is how easily those collisions can occur or be deliberately found.

For detecting accidental damage, simple checksums are genuinely useful: send data across a network along with a value calculated from it, and if the receiving system's own calculation doesn't match, something went wrong. One especially common family is the CRC — Cyclic Redundancy Check — designed to efficiently catch common patterns of accidental corruption in storage, networking, and compressed files. If you've ever opened a damaged ZIP archive and gotten a CRC error, you've seen this exposed directly: the software calculated what it expected, and the numbers didn't agree.

But accidental corruption and deliberate tampering are different problems. If an attacker knows which checksum algorithm you're using and that algorithm is weak, they may be able to modify a malicious file until it produces the same checksum as the legitimate one — and now your integrity check happily reports that everything matches. Security needs something much harder to manipulate on purpose. That's where cryptographic hash functions come in.

A Fixed-Size Fingerprint for Any Amount of Data

A cryptographic hash function accepts data of essentially arbitrary size — a single word, a photograph, an entire operating-system image — and produces a fixed-size result called a hash, hash value, or digest. For a given algorithm, the output length never changes regardless of input size. SHA-256 always produces a 256-bit hash, whether the input is five bytes or five billion.

Since computers display binary awkwardly, hashes are commonly shown in hexadecimal — sixteen symbols, 0 through 9 and A through F, with each digit representing four bits. A 256-bit hash therefore needs exactly 64 hexadecimal characters, which is why SHA-256 output has that distinctive long, blocky look.

A useful hash function is deterministic: the same input always produces the same output. That's what makes verification possible — a developer publishes a SHA-256 value, you calculate your downloaded file's hash yourself, and if the two match exactly, you have strong evidence your file contains precisely the bytes used to produce that published hash. And a good cryptographic hash exhibits the avalanche effect: change even one bit of the input and the output changes dramatically and unpredictably, not just slightly. That property matters enormously for integrity checking — you don't want a slightly modified malicious file producing a hash that visibly resembles the legitimate one.

It's worth being precise about what a hash is not. It isn't compression — you can't recover a 10 GB file from its 32-byte hash, because the hash deliberately throws away almost all the original information. And it isn't encryption, which is designed to be reversible with the right key. A hash has no "decrypt" operation that reconstructs the original input; it's intentionally one-way. Calling something "encrypted with SHA-256" is technically a category error — SHA-256 is a hash function, not an encryption algorithm.

One-Way Doesn't Mean Unguessable

"One-way" means there's no practical mathematical shortcut to run the function backward. It doesn't mean nobody can ever figure out what produced a given hash — you can always try candidate inputs, hash each one, and check for a match. If the original input comes from a small, predictable set of possibilities, guessing can still be fast, hash function or no hash function.

This is exactly why storing passwords is trickier than it first appears. Simply hashing a password with a fast, general-purpose algorithm like SHA-256 isn't considered sufficient on its own, because SHA-256 is too fast — modern hardware can compute billions of hashes per second, so an attacker who steals a password database can simply guess common passwords, hash each guess, and compare against the stolen values at enormous speed. For password storage, you actually want something slow and expensive. That's what dedicated password-hashing algorithms like Argon2, scrypt, bcrypt, and PBKDF2 are built for, as OWASP's own password storage guidance describes — deliberately requiring extra computation or memory so that a legitimate login barely notices the cost, while an attacker trying billions of guesses pays that cost billions of times over.

Password systems also add a salt — random data combined with the password before hashing, so that two users with the identical password end up with completely different stored values, and so precomputed lookup tables (rainbow tables) built against common passwords become useless, since an attacker can no longer reuse one table against every account.

Verifying Downloads — and Where That Trust Can Break

Return to the operating-system image: the developer publishes a SHA-256 value, you calculate the same hash over your downloaded file, and the values match. That tells you the file you have matches the data used to produce that published hash. But there's a catch — you also have to trust the hash you were given in the first place. If an attacker compromises the download server and replaces both the installer and the published hash displayed beside it, your calculation will match perfectly. The hash function worked exactly as designed. Your trust model didn't.

This is where digital signatures come in. A developer with a private cryptographic key can sign information tied to the software, and anyone can verify that signature with the corresponding public key — establishing not just "does this match a reference value" but "was this reference value authorized by whoever holds this specific private key, and has it been altered since." Rather than signing an entire multi-gigabyte file directly (an expensive operation), systems sign the compact hash instead; because changing the file changes its hash, protecting the hash effectively protects the whole file. A tiny fixed-size digest standing in for an enormous amount of data during other cryptographic operations is one of hashing's most elegant uses.

How Large Is 256 Bits, Really?

A 256-bit value has 2²⁵⁶ possible combinations — roughly 1.16 × 10⁷⁷, a number in the same neighborhood as estimates for the number of atoms in the observable universe. Randomly stumbling onto one specific 256-bit value is effectively impossible. There's a subtlety worth knowing, though: finding any two inputs that happen to share a hash is a different, easier problem than matching one specific existing hash, related to the birthday paradox — for an n-bit hash, that kind of generic collision search takes roughly 2^(n/2) work rather than 2^n. For SHA-256, that's still around 2¹²⁸ operations, a number so large it remains firmly out of reach of any practical attack. NIST's own Secure Hash Standard, FIPS 180-4, formally defines SHA-256 and the rest of the SHA-2 family specifically to provide this kind of collision resistance for integrity verification, digital signatures, and authentication systems.

Not every hash function keeps these properties forever, though. MD5 and SHA-1 were both once widely used for exactly this kind of security work, and both have since had practical collision attacks demonstrated against them. That doesn't mean they instantly became useless for catching an accidentally flipped bit — it means their guarantees against a deliberate attacker no longer hold, which is why "it uses a hash" was never really the right question. The right question is always which hash, and what it's being asked to protect against.

Hashing Shows Up Everywhere Once You Notice It

Because hash algorithms process data incrementally — reading one block, updating internal state, reading the next — they never need an entire file loaded into memory at once, which makes them practical even for enormous files and continuous data streams. Git leans on this directly: it builds identifiers derived from commits, trees, and file contents, so changing content anywhere changes the resulting identifier, letting Git track relationships between versions and detect exactly when something differs. Git's own documentation on its internal object model describes this content-addressed approach explicitly — rather than asking "where is the file named X," the system can effectively ask "give me the data whose digest is Y," which turns out to be useful well beyond version control, including deduplication, caching, and distributed storage systems. Blockchains made terms like SHA-256 familiar to a much wider audience by using existing cryptographic hashes to link data structures together, though the hash function itself long predates cryptocurrency and was never built for that purpose specifically — the same mathematical tool just turned out to be useful there too.

Hashing isn't automatically anonymizing, either. If a database replaces email addresses with their SHA-256 hashes, an attacker can simply hash likely email addresses and compare results — the hash function is one-way, but the input space was guessable to begin with. That's the same underlying issue as weak password hashing, just applied somewhere people don't always expect it.

The Bard's Take

A hash is one of those computer-science ideas that seems almost too simple to deserve much attention: take data, perform a calculation, get a smaller value, compare it later. But that basic idea becomes extraordinarily powerful once the algorithm has the right mathematical properties. A simple checksum can tell you whether data was probably damaged by accident. A cryptographic hash like SHA-256 can tell you something much stronger — whether data matches even when someone may be actively trying to fool you — by taking practically any amount of information and reducing it to a fixed 256-bit digest that changes completely the moment anything in the input changes.

It isn't encryption, because there's no key that reverses it and reconstructs the original file. It isn't compression, because almost all the original information is deliberately discarded. It's a mathematical fingerprint: compact, deterministic, and fiercely sensitive to change. That fingerprint can confirm a downloaded installer matches what a developer actually published, let a digital signature protect a massive file without expensive operations on every byte, identify exact duplicate files regardless of their names, and give Git a way to track an entire project's history.

A hash does have real limits, though. It doesn't prove its own source is trustworthy — a compromised server can swap both the file and the hash together. It doesn't make a predictable input unguessable. And an aging algorithm doesn't become safe again just because its output still looks impressively complicated. The useful question was never "are we using a hash," but what problem you're actually trying to solve — and when that problem is "are these two enormous piles of data exactly the same," sixty-four little hexadecimal characters can answer it with remarkable confidence.

Sources