Contents
- Introduction
- ECB/CBC mode — concept
- Getting started
- Stage 1 — Identifying the block size
- Stage 2 — Identifying whether it’s CBC or ECB
- Stage 3 — Where the input begins
- Stage 4 — Identifying the suffix after the input
- Exploitation — hands-on
- Conclusions
- What I learned from this study
Introduction
Natas28 is one of the levels of Natas, the web security wargame from OverTheWire (overthewire.org), focused on common vulnerabilities in web applications.
In this level, the application has a search field (“query”) whose content gets encrypted and sent back as a URL parameter (Base64-encoded), with no access to the source code.
The challenge exploits a weakness in encryption that uses ECB-mode block ciphers deterministically, with 16-byte blocks. Because ECB mode encrypts each 16-byte block separately and always the same way for the same plaintext, you can manipulate the length of the input to isolate and “reassemble” ciphertext blocks, effectively changing what the server decrypts without ever knowing the key.
Since the content being encrypted is a SQL query, in this specific case it’s possible to escalate the cryptographic weakness into a full SQL injection, even without access to the encryption keys.
In this scenario, the plaintext typed into the input is sent to the server, where it’s encrypted on the backend and returned as a redirect to another endpoint carrying the encrypted result, like this:
http://natas28.natas.labs.overthewire.org/search.php/?query=[encrypted-value]
We have no access to the encryption routine, but in this scenario we can freely observe the values we put into the input and the ciphertext the server produces in response.
The goal is to manipulate the encryption so we can insert or remove information from the protected content the server will process, breaking the method’s trust model and injecting malicious content.
My intent here is not to hand you the challenge’s solution, but to document the process of understanding, reconnaissance and possible exploitation, since a similar scenario is perfectly plausible in real-world systems.
I should be upfront that I’m not a cryptography specialist. I know the basics I’ve used and implemented as a developer. So what follows is a perspective rooted in that starting point — the same one, I imagine, that most working developers share.
Given my lack of specialization, I’ve taken risks copying and pasting encryption routines from forums like StackOverflow, without fully understanding the various parameters you can set when building that kind of routine — a case in point: ECB or CBC mode.
📌 Concept — the next two sections (ECB/CBC mode) are theoretical. If you already know the difference between the two modes, feel free to skip straight to Getting started.
ECB/CBC mode
In ECB mode, the message is split into fixed-size blocks (usually 16 bytes for AES) and each block is encrypted completely independently, using the same key and algorithm, with no relationship between blocks. That means identical plaintext blocks always produce identical ciphertext blocks. That property is exactly what makes ECB insecure for most applications: an attacker can spot repeating patterns in the ciphertext, reorder blocks, or even swap in known ciphertext blocks to manipulate the decrypted content. For that reason, ECB is considered the weakest of the block cipher modes of operation.
In CBC mode, each plaintext block is XORed with the previous ciphertext block before being encrypted, creating a chain of dependencies across every block in the message. Since the first block has no “previous block,” an initialization vector (IV) is used — a random value generated for each encryption operation — which gets XORed with the first plaintext block. This dependency solves ECB’s main problem: identical plaintext blocks produce different ciphertext blocks (as long as the IV differs or the preceding content changes), eliminating the visible patterns. On the other hand, CBC brings its own classic weaknesses, such as vulnerability to padding oracle attacks (when the server reveals, even indirectly, whether the padding of a decrypted message is valid) and the need for a truly random, unpredictable IV to stay secure.
I’ll focus on ECB here, but the idea carries over to CBC too, which adds a layer of attempted difficulty through the XOR step — though it’s common knowledge that XOR is trivially reversible.
From here on the post moves into the hands-on part: first reconnaissance of how the encryption behaves, then the actual exploitation.
Getting started
With that introduction out of the way, let’s move into reconnaissance and the possible exploitation of deterministic ECB encryption. Reconnaissance breaks down into 4 stages.
In this specific scenario, since the encrypted content is sent in the open as part of the URL (http://natas28.natas.labs.overthewire.org/search.php/?query=[encrypted-value]), you can change the value to some random invalid one, triggering an expected exception in the decryption method. Looking at the returned error message, we get useful information such as:
1 | Zero padding found instead of PKCS#7 padding |
This kind of exception is usually a strong signal that the method is using CBC or ECB mode. This is the first step of reconnaissance and exploitation.
Stage 1 — Identifying the block size
TL;DR: by varying the input length and watching when the ciphertext grows by a whole block, we find the block size is 16 bytes.
The next step is identifying the block size the routine uses — whether it’s configured for 16 bytes, 32 bytes, etc. Since in this scenario we can submit any value to generate the ciphertext, we can test with an arbitrary input and observe the byte length of the resulting ciphertext. Then we keep growing the input one character at a time until we see the ciphertext’s block count increase, at which point we have the exact block size in use.
In practice, this can be done as follows:
- Search with an empty string; the resulting base64 is:
1 | Base64: |
- Search with a single character:
1 | Base64: |
The same byte count is produced, because the final size is always a multiple of the block size, even when there are spare bytes left over to fill out a block — that’s where padding kicks in.
This same byte count keeps showing up until we hit the limit of 12 characters; sending 13 characters instead, we see a whole new block get added. So:
- Search with ‘13’ characters (e.g. aaaaaaaaaaaaa); the resulting base64 is:
1 | Base64: |
Adding the 13th character makes the output jump by 16 bytes (32 extra characters, each pair being 1 byte). That means each block is 16 bytes — the first fact the solver.py script relies on.
To do these conversions, you can use https://cyberchef.org and chain ‘URL Decode’ / ‘From Base64’ / ‘To Hex’ to inspect the data, or use the blocks.py script, for example.
Stage 2 — Identifying whether it’s CBC or ECB
TL;DR: if ciphertext blocks repeat for repeated inputs, it’s ECB. In CBC that can’t happen, since every block depends on the one before it.
The identification process can be done like this:
- Submit a long string of identical characters, e.g. 32 to 64 ‘A’s.
- Decode it and split it into 16-byte chunks (the block size we found in the previous stage).
- Look for two identical blocks.
Put simply, in ECB an identical input block always produces an identical output block (each block is encrypted in isolation, with no chaining). If repeated ciphertext blocks show up, it’s ECB.
Using blocks.py, which generates and organizes the 16-byte blocks for different encrypted payloads, we can see that blocks repeat even across different inputs.
1 | [empty string] -> 80 bytes, 5.00 blocks of 16B |
In this analysis, the tail blocks show up REPEATED at different positions, even with different blocks BEFORE them. In CBC that would be impossible. So: ECB.
Example:
Block: a77e8ed1aabe0b5d05c4ffe6ac1423ab
Block: 478eb1a1fe261a2c6c15061109b3feda
Stage 3 — Where the input begins
TL;DR: by repeating known “twin” blocks while varying a test prefix, we can calculate that there’s a fixed 38-byte prefix before our input (the text
SELECT * FROM data WHERE query LIKE '%).
This is a case where we can tell there’s always a fixed prefix in the generated ciphertext, and that’s specific to this particular study — in other scenarios there may or may not be a prefix, depending entirely on the logic of the system under review.
We notice this because, no matter what input we submit, blocks 0 and 1 always stay the same, meaning our input content must start somewhere right around the beginning of block 2.
To manipulate the blocks, we need to know exactly where our input starts inside the generated ciphertext, so the goal is to figure out the size of that prefix.
The technique here is to force and look for twin blocks by varying a run of identical characters (‘A’s, for example).
This works by looping over the number of ‘A’s placed in front of 32 ‘Q’s. The number 32 is chosen because we want 2 full blocks, and each block is 16 bytes. That quantity would change if the block size were something other than 16 bytes.
The sequence looks like:
- “QQQ…Q” (0 A’s + 32 Q’s)
- “A” + “QQQ…Q” (1 A + 32 Q’s)
- “AA” + “QQQ…Q” (2 A’s + 32 Q’s) … and so on.
For each loop iteration:
- Decode, split into 16-byte blocks, and look for two identical, adjacent blocks in hex.
- On the first iteration where twin blocks show up, note down:
- f = how many ‘A’s were inserted
- j = the index where the twin blocks start
Which gives us: prefix = j*16 - f
Here’s the result one step before the match, and then with the actual match found:
1 | (9 A's + 32 Q's) |
The logic is: where the blocks repeat (3 and 4) we know that stretch is made up entirely of ‘Q‘ characters, and what comes right before it is ‘A‘.
Since we know exactly how many ‘A’s we inserted, we can calculate the byte where our input actually starts, and from there the size of our prefix.
Reading it:
- With f=10 ‘A’s, blocks 3 and 4 came out IDENTICAL.
- prefix = 3*16 - 10 = 48 - 10 = 38 bytes.
- This matches the plaintext “
SELECT * FROM data WHERE query LIKE '%“, which is exactly 38 characters long.
It would be equivalent to say the prefix runs up through this point inside block 2:
block 0: 1be82511a7ba5bfd578c0eef466db59c
block 1: dc84728fdcf89d93751d10a7c75c8cf2
block 2: 5f22a727f625[...]
So we can confirm the prefix is 38 bytes long, and the input begins right after it.
Stage 4 — Identifying the suffix after the input
TL;DR: using a classic ECB byte-at-a-time attack, we recover the suffix
%'that closes the SQL statement right after our input.
In this specific scenario, besides the prefix, we can also tell there’s a suffix appended to the final result. Since we suspect the final content is a SQL query, identifying that suffix is key to being able to manipulate the statement.
Using what we’ve learned so far, we can now recover the suffix by running a byte-by-byte recovery test — a technique known as ECB’s “byte-at-a-time“ attack.
The process starts by figuring out the LENGTH of the suffix, growing the input until the ciphertext gains +1 block. The difference tells us how many suffix bytes exist.
Then, for each position in the suffix:
- Build a “short frame” that lands the target suffix byte as the LAST byte of a known block.
- Store that target ciphertext block.
- Try all 256 possible values (short + known + guess) until the ciphertext block matches. The guess that matches is the suffix byte.
- Slide forward and repeat for the next byte.
In this specific scenario, the routine escapes single quotes to try to prevent SQLi, so inserting ' turns the content into \'. That matters, because no tested byte can ever reproduce a literal quote at that position — but that itself lets us assume the suffix ends in a quote.
The manual walkthrough would go like this:
Assuming prefix_pad = 10 (from the previous stages), to find the 1st suffix byte:
- Submit
"A"*10 + "A"*15(i.e., align things and leave 15 A’s in the target block).
This pushes the 1st suffix byte into the LAST position of that block.
Note the target ciphertext block (the one at indexstart). - Submit
"A"*10 + "A"*15 + X, varying X across the printable
characters' ', '%', '&', ...For each X, compare the SAME block.
When the block matches the one from step 1, that X is the 1st suffix byte. - Expected result: the 1st suffix byte is ‘
%‘. The 2nd would be the quote,
which will never match, since it’s being escaped.
Suffix conclusion: %'
Example input/output:
- prefix=38
- prefix_pad=10
In practice it looks like this:
1 | "A"*10 + "A"*15 (TEMPLATE) |
The solver.py script automates this entire process.
With reconnaissance done (we know the block size, that it’s ECB, the 38-byte prefix and the
%'suffix), we move on to the actual exploitation.
Exploitation
With all this information in hand, we can now move into exploitation and manipulate the ciphertext, inserting or removing information from the content that the server-side routine will successfully decrypt — trusting the content it receives as valid and intact.
Since in ECB each block is an independent “box“, we can throw one away by deleting it without affecting the rest. The idea is to “push“ the escape character, or the closing of the SQL statement, into a block we’re going to discard, and use that to manipulate the SQL.
So the plan is to work with 15 “pusher“ characters: "AAAAAAAAAABBBBBBBBBBBBBBB" + ' + <injection>
The server will escape the quote ('), turning it into \'. That backslash gets pushed into the next block, and we discard the block that ends up as "AAAAAAAAAABBBBBBBBBBBBBBB\“
Step 1: Trying to force a SQL error
TL;DR: the test payload returns no error at all, which reveals the application is suppressing SQL errors. We’ll have to proceed in blind mode.
The idea is to break the SQL syntax to see what comes back.
Payload that breaks the SQL: AAAAAAAAAAAAAAAAAAAAAAAA (10 ‘A’s to complete the first block, plus 14 ‘A’s so the %’ closing characters of the SQL land in the same block)
Using blocks.py, we get these blocks:
1 | INVALID-PAYLOAD -> 96 bytes, 6.00 blocks of 16B |
Then, using assemble.py, we reassemble the new payload:
1 | G%2BglEae6W%2F1XjA7vRm21nNyEco%2Fc%2BJ2TdR0Qp8dcjPJfIqcn9iVBmkZvmvU4kfmyiW3pCIT4YQixZ%2Fi0rqXXY5FyMgUUg%2BaORY%2FQZhZ7MKM%3D |
Unfortunately, no error comes back, which indicates that SQL errors are being suppressed, so we need to proceed in ‘blind‘ mode after all.
Step 2: Testing for SQL injection
TL;DR: the payload
UNION SELECT 1injects successfully, and the “1” shows up reflected in the HTML, confirming the SQL injection works.
The idea is to use the payload “AAAAAAAAAABBBBBBBBBBBBBBB' UNION SELECT 1,2 #“ to kick off the injection. Whatever lands in the block of ‘B’s gets discarded. The plan is to send 15 'B's + '; since we expect the quote to get escaped with \, internally the payload becomes “AAAAAAAAAABBBBBBBBBBBBBBB\' UNION SELECT 1,2 #“, which makes block 3 end up as “BBBBBBBBBBBBBBB\“ — and once we drop that block, the payload effectively becomes “AAAAAAAAAA' UNION SELECT 1,2 #“
Running it through blocks.py gives us:
1 | AAAAAAAAAABBBBBBBBBBBBBBB' UNION SELECT 1,2 # -> 128 bytes, 8.00 blocks of 16B |
Using the assemble.py script to rebuild the encrypted value without block 3 (a090b8345b6d2823ac11042c72490053), we get:
1 | Hex: |
The HTML comes back blank with this payload, since SQL errors are being suppressed. When the UNION executes successfully, it will add the UNION‘s result to whatever result would normally come back.
So switching the payload to use 1 column (AAAAAAAAAABBBBBBBBBBBBBBB' UNION SELECT 1 #), splitting it into blocks, dropping block 3, reassembling it and generating a new URL-encoded base64 to send in the request, we get this result:
<h2> Whack Computer Joke Database</h2><ul><li>1</li></ul>
The “1” in the HTML is exactly the result of the UNION SELECT 1 we expected.
From here, it’s just a matter of following through with standard SQL injection techniques to extract database data such as table and column names, which we’ll use to pull the content we’re after.
Step 3: Extracting the column names of the ‘users’ table
TL;DR: a
UNION SELECTagainstinformation_schema.columnsreveals the column names:usernameandpassword.
Following the same approach, we use the payload:
“ UNION SELECT GROUP_CONCAT(column_name SEPARATOR 0x0a) FROM information_schema.columns WHERE table_name=0x7573657273“
Keep in mind 0x7573657273 is the hex equivalent of ‘users‘
This injection returns the results:
- password
- username
Step 4: Extracting the content of the ‘users’ table
TL;DR: with the column names in hand, a second
UNION SELECTextracts the username/password pair, which is the challenge’s flag.
Following the same idea, we can now run a SELECT against the ‘users‘ table using the fields obtained in the previous step:
“ UNION SELECT GROUP_CONCAT(username,0x3a,password SEPARATOR 0x0a) FROM users“
This injection returns the challenge’s flag.
All the steps above are automated by the following scripts:
Conclusions
Any block cipher running in ECB mode inherits this weakness, regardless of the underlying algorithm. AES-ECB, 3DES-ECB, Camellia-ECB, Blowfish-ECB — they all suffer from the same problem, because the weakness lies in how the blocks are combined, not in the cipher itself.
CBC chains blocks together via XOR with the previous ciphertext block. That breaks both the determinism (given a random IV) and the independence, so the kind of copy-paste-and-reassemble attack done here simply doesn’t work against it. But CBC has its own weaknesses — bit-flipping and the padding oracle attack mentioned earlier. Put another way, CBC is vulnerable in a different way, especially when the system reveals a padding (PKCS#7)-type error, since a padding oracle attack depends on the system revealing whether the padding ended up valid or not.
What I learned from this study
- Block cipher modes of operation: the practical difference between ECB (each block encrypted independently and deterministically) and CBC (blocks chained via XOR, dependent on an IV).
- Why ECB is insecure: identical plaintext blocks produce identical ciphertext blocks, creating patterns that are visible and exploitable without ever knowing the key.
- CBC’s weaknesses: vulnerability to padding oracle attacks and the importance of a truly random IV.
- Reverse-engineering a “black-box” encryption routine: how to infer a system’s behavior by observing only inputs and outputs, with no access to source code or the key.
- How to identify the block size a cipher uses, by varying the input length and watching for jumps in ciphertext size.
- How to tell ECB apart from CBC in practice: looking for repeated ciphertext blocks from repeated inputs.
- The “twin blocks” technique for figuring out the exact size of a fixed prefix that precedes the user’s input inside the encrypted content.
- The classic ECB byte-at-a-time attack: how to recover, byte by byte, an unknown suffix that follows input the attacker controls.
- Manipulating ECB blocks (“block-box” manipulation): how to discard an entire block (for example, one containing an escape character) to change what the server decrypts, without ever knowing the key.
- Escalating a cryptographic weakness into SQL injection: turning a cipher-mode vulnerability into a working SQL injection.
- Blind SQL injection exploitation techniques: confirming an injection via
UNION SELECTeven when the application suppresses SQL errors. - Using hexadecimal literals to bypass escaping filters: representing strings like
0x7573657273(equivalent to'users') avoids quotes in the payload altogether, bypassing routines that escape or block the'character as an anti-SQLi protection. - Extracting metadata and data via SQLi: using
information_schema.columnsto discover column names, then extracting the data from a table. - Practical tools and workflow: combining custom scripts (
blocks.py,assemble.py,solver.py) with external tools (CyberChef) to decode, analyze and reassemble encrypted payloads. - Pentest/CTF mindset: documenting the reconnaissance and exploitation process in a structured way, useful both for wargames (Natas/OverTheWire) and for real vulnerabilities in production systems.