Even if you've never seen assembly, you can start here. From concepts to solving a real function line by line — trying every step inside tlk-hex.
Disassembler, assembly, registers, memory — what are we dealing with?
PE/ELF headers, sections, imports, entry point.
Registers, common instructions, calling convention, the stack.
Let's solve a function with tlk-hex, start to finish.
Reading if / loop / switch from the graph.
Where and how to practise; legitimate exercise sources.
Combine everything in one real workflow: string → xref → decision → verify.
Change behaviour instead of finding the key: flip a branch, skip a check.
When the code is hidden: entropy, packer signatures, imphash and sparse imports.
You write a program, the compiler turns it into machine code: the bytes the processor runs directly. The source code is usually not available — all we have is those bytes. Reverse engineering is working backwards from those bytes to understand what the program does.
The processor reads bytes and runs them as instructions. For example the bytes
48 83 EC 28 actually mean "reserve 0x28 bytes on the stack". A disassembler turns
those bytes into human-readable assembly:
48 83 EC 28 sub rsp, 28h ; reserve 40 bytes on the stack
tlk-hex decodes the whole file this way, then ties the functions, calls and data together.
rax, rbx, rcx…0x140001008.mov, add, call…call, returns with ret.00–FF
fits exactly into two hex digits. That's why addresses and values are always written in hex. In tlk-hex,
H converts a number to decimal.An .exe isn't just code. For Windows to load it into memory correctly it has a layout:
the PE (Portable Executable) format. On Linux the equivalent is ELF.
The file is split into sections by purpose. In tlk-hex, Shift+F7 shows them all:
| Section | What's inside |
|---|---|
| .text | Executable code. Functions live here. |
| .rdata | Read-only data: strings, import table, constants. |
| .data | Writable global variables. |
| .rsrc | Resources: icons, version info, dialogs. |
| .pdata | x64 exception/unwind info — helps find function boundaries. |
CreateFileW, send…). The strongest clue to what it does.0x140000000 for most x64 EXEs).It looks scary but a few dozen instructions cover 90% of the work. First, the registers:
| 64-bit | 32-bit | Typical use |
|---|---|---|
| rax | eax | Return value; general math |
| rcx rdx r8 r9 | ecx edx | First 4 arguments to a function (x64) |
| rbx rsi rdi | ebx esi edi | Preserved general variables |
| rsp | esp | Stack pointer |
| rbp | ebp | Stack frame base |
| rip | eip | Address of the next instruction |
Smaller parts of a register are also used: rax → eax (32) →
ax (16) → al (8 bit).
| Instruction | Meaning |
|---|---|
| mov a, b | a = b — copy a value |
| lea a, [b] | a = address of b — compute an address (doesn't always read memory) |
| add / sub | Add / subtract |
| xor a, a | Zero a register (XOR with itself = 0) |
| cmp a, b | Compare (a - b) and set flags |
| test a, a | Check whether a is zero (very common) |
| jmp | Unconditional jump |
| jz / je | Jump if equal / zero · jnz/jne: if not |
| call | Call a function |
| ret | Return from a function |
| push / pop | Push onto / pop off the stack |
When a function is called, arguments go in order into rcx, rdx, r8, r9; extras go on the
stack. The return value is in rax. So when you see this:
mov rcx, 1 lea rdx, aHello ; "hello" call WriteMessage ; WriteMessage(1, "hello") mov ebx, eax ; save the return value
…you're actually reading the call WriteMessage(1, "hello").
Local variables live on the stack. tlk-hex names them automatically as var_18 (local)
and arg_8 (argument); the text [rsp+88h+var_18] means "a local
variable of this function". Hover over it and press N to give it a meaningful name.
Now let's read a real function. Take the example from the home page and solve it line by line:
140001008 ParseConfigFile proc near 140001008 mov r11, rsp 14000100B sub rsp, 88h ; reserve stack space 140001012 mov rax, cs:__security_cookie ; stack protection 140001019 xor rax, rsp 14000101C mov [rsp+88h+var_18], rax 140001053 lea rcx, aConfigLoaded ; "config loaded: %s" 14000105A mov edx, esi 140001073 call LogMessage ; LogMessage("config loaded: %s", esi) 140001080 call cs:__security_check_cookie 140001085 add rsp, 88h ; give the space back 14000108C retn
What did we read?
sub rsp + security_cookie): the standard opening of every compiler-generated function. Protection against stack overflows. The logic isn't here.lea rcx, aConfigLoaded takes the address of a string, call LogMessage prints it. So this function logs something.security_check + add rsp + retn): the standard closing.So once you strip the noise, the essence of the function is one line: log the "config loaded" message. Half of real analysis is recognising these standard patterns and focusing on what matters.
call target to step
into that function.socket/connect and there's networking; see CryptEncrypt and there's encryption.Press Space for the graph. The arrows between blocks are the program's logic. Learn a few patterns:
A cmp + conditional jump is a branch. Two arrows leave the block:
green = condition true, red = false.
cmp eax, 0 jz loc_error ; if eax == 0 jump to error ; ... falling through here means eax != 0 (success path)
In C: if (eax == 0) goto error;
When you see an arrow in the graph that goes back up, that's a loop: the block jumps back to an earlier
one. It usually ends with a counter (inc/dec) and a cmp.
Branching many ways based on a value. tlk-hex resolves jump tables automatically and adds comments like
jumptable ... case 3 — you read directly which value goes where.
if/goto for you. Not exact, but it conveys
the logic fast.Reading doesn't teach; solving does. Safe and legitimate ways to begin:
So far we've seen the pieces separately: strings, imports, xrefs, assembly, graph and pseudocode. Now we'll combine them into one real analysis workflow. Don't rush — the goal of this module isn't speed, it's learning how to ask a binary questions.
We'll pretend we have no source code and answer these only from the file itself:
First we write the program we'll analyse. Create a file sample.c:
#include <stdio.h> #include <string.h> static int classify(const char *mode) { if (strcmp(mode, "debug") == 0) return 2; if (strcmp(mode, "safe") == 0) return 1; return 0; } int main(int argc, char **argv) { const char *mode = argc > 1 ? argv[1] : "safe"; int result = classify(mode); printf("mode=%s result=%d\n", mode, result); return result; }
For now, forget what it does. We only use the source to create our target; once the EXE exists we look at it purely as a binary. Compile it (whichever you have):
cl /O2 /Fe:sample.exe sample.c
gcc -O2 -o sample.exe sample.c
We now have sample.exe. From here on, close the source.
Drag sample.exe onto tlk-hex. Auto-analysis starts and recovers sections, functions,
calls, strings, xrefs and imports.
When it finishes, don't dive into assembly yet. Get a feel for the file first. In reverse engineering the first goal isn't detail, it's context — reading instructions one by one without knowing where you are just wears you out.
Open Strings with Shift+F12. You'll look for text like:
debug safe mode=%s result=%d
These three lines already tell us a lot before reading a single instruction. debug and
safe are probably values compared against user input; mode=%s
result=%d suggests the program prints a result at the end.
Select the debug string and press X. This shows where it's used
(cross-references). Jump to a code reference. Do the same for safe.
You'll see both strings' references land in the same function or very close code. A strong clue:
"debug" and
"safe".Now look at the functions the program borrows from outside. Names vary a little by compiler, but you'll see:
| Import | What it does |
|---|---|
| strcmp | Compares two C strings (returns 0 if equal) |
| printf | Prints formatted text |
Now three independent pieces of evidence have piled up: "debug",
"safe" and strcmp. Together they point to a strong likelihood:
the program compares a string against "debug" and "safe". We still won't be sure until we see it in
assembly.
Go to the start of the function you reached from the debug xref. Its auto name might be
sub_140001000 — we don't know its real name yet. Inside you'll see a pattern like this
(exact instructions vary by compiler; what matters is reading the pattern):
lea rdx, aDebug ; 2nd arg: "debug" mov rcx, rbx ; 1st arg: the input call strcmp ; strcmp(input, "debug") test eax, eax ; is the result 0? (equal?) jz loc_debug ; if equal, jump to the debug block
The first two lines prepare strcmp's arguments. strcmp returns
0 when the strings are equal; test eax, eax checks whether the result is zero and
jz (jump if zero) goes to that block. So the logic is clear:
if (strcmp(input, "debug") == 0)
We roughly know what it does now. On the function, press N and rename
sub_140001000 to classify_mode.
Naming isn't a luxury, it's the heart of the method. At first you see dozens of
sub_140001000; as you turn them into classify_mode,
print_result, load_config and so on, the file gets more readable
every pass.
Inside the function, press Space; instead of text you get the control-flow graph. For this example we roughly expect this decision tree:
input
│
▼
input == "debug"?
┌──────┴──────┐
yes no
│ │
▼ ▼
return 2 input == "safe"?
┌─────┴─────┐
yes no
│ │
▼ ▼
return 1 return 0What looks tangled in assembly reads at a glance in the graph. Follow the no (false) path of the first
comparison and after a while you'll find a second comparison using the "safe"
string — same logic:
lea rdx, aSafe mov rcx, rbx call strcmp test eax, eax jz loc_safe ; if (strcmp(input, "safe") == 0)
Now check which value each path returns. On x64 Windows the integer return value comes back in
eax:
mov eax, 2 ret ; return 2; (debug path) mov eax, 1 ret ; return 1; (safe path) xor eax, eax ret ; return 0; (xor = zero it; default path)
xor eax, eax zeroes the register — i.e. return 0;. Recognising this
little pattern pays off; you'll see it everywhere.
Combining the evidence, the binary is telling us:
int classify_mode(const char *input) { if (strcmp(input, "debug") == 0) return 2; if (strcmp(input, "safe") == 0) return 1; return 0; }
We recovered the function's behaviour without the source — and that's exactly the point of reverse engineering: not getting the source back verbatim, but understanding the behaviour well enough.
Now press F5; tlk-hex turns the same function into C-like pseudocode and shows something very close to what you worked out by hand. It's fast, but a warning:
On classify_mode, press X again — this time you see who calls the
function. You'll probably land in main or near it:
mov rcx, rbx ; input call classify_mode mov esi, eax ; save the return value (used later)
Then back to Strings, follow mode=%s result=%d with X; nearby you'll see a
printf call. So we've recovered the whole chain:
command-line input → classify_mode → 0 / 1 / 2 → printf → exit
Make your findings permanent:
sub_140001000 → classify_mode,
aDebug → mode_debug, aSafe → mode_safe…Reverse engineering isn't only reading a binary; it's recording what you find in an orderly way. A well-named database looks completely different — far more readable — a few hours in.
A good session ends with a short conclusion. For this example:
sample.exe takes a "mode" value from the command line. mode == "debug" → result = 2 mode == "safe" → result = 1 any other value → result = 0 The result is printed with printf as "mode=%s result=%d".
We recovered the program's core behaviour without ever looking at the source. That was the whole point.
This loop works on almost any file — no need to memorize it, it becomes second nature after a few runs:
1. Open the file, wait for auto-analysis 2. Look at sections / imports / strings (context) 3. Pick an interesting string or import 4. Go to code with xref (X) 5. Study the function's graph (Space) 6. Follow the call targets 7. Extract the conditions and return values 8. Quick check with F5 pseudocode 9. Verify in assembly 10. Name the functions (N) 11. Add comments (;) 12. Save (Ctrl+W) 13. Write the result in a few sentences
Strings are the strongest clues. If you see these, stop and follow
string → X → function:
error failed password token config login connect http registry file debug
Imports reveal which capabilities the program uses:
| Import group | What it points to |
|---|---|
| CreateFileW · ReadFile · WriteFile | File operations |
| RegOpenKeyExW · RegSetValueExW | Registry usage |
| connect · send · recv | Networking |
| CryptEncrypt · CryptDecrypt | Encryption |
tlk-hex currently does mostly static analysis — it inspects the program without running it. So assembly analysis, function discovery, xrefs, string/import analysis, the graph, pseudocode, PDB symbols, hex inspection and patch preparation all work comfortably. But some things can't always be seen with pure static analysis:
values computed at runtime encrypted code decrypted in memory self-modifying code runtime unpacking (packers that unpack while running)
In those cases dynamic tools like a debugger come in. Still, most of the work starts static; you use static analysis to find what to look for, then switch to dynamic if needed.
Recompile the sample with a small change (try to analyse it without looking at the source again):
if (strcmp(mode, "admin") == 0) return 3;
Open the new EXE in tlk-hex and find, from the binary only:
If you can find all of these from the binary, you've completed your first real reverse-engineering workflow.
Sometimes you don't want to find the right password/serial — you want to change the program's behaviour: e.g. flip a check so it "always passes". That's patching: editing the bytes on disk. tlk-hex does this safely — it never touches the original file and keeps changes separately.
Most checks end in a cmp/test + a conditional jump. A typical
"deny on failure" pattern:
cmp eax, 0DEADBEEFh jnz loc_denied ; if NOT equal, jump to deny ; ... falling through here is the "granted" path
There are a few ways to change the flow — each is a one-byte edit:
| Goal | How | Byte |
|---|---|---|
| Invert the condition | jnz → jz (or vice versa) | 75 ↔ 74 |
| Never take the jump | NOP the conditional jump | 90 90 |
| Always jump | Make it unconditional: jnz → jmp | 75 → EB |
74 = jz/je, 75 = jnz/jne,
90 = nop, EB = short jmp. Memorizing these
four bytes speeds you up a lot.jnz loc_denied).75 to 74 (or whatever you need). F2 / Esc to leave.jz — the logic is inverted.The patch lives in the database for now. To produce a permanent file:
Two of the practice crackmes are made for this module:
0DEADBEEFh). Find the branch
after the cmp and patch it → "Access granted".Grab both from the crackmes release. The other levels fit the workflow in Module 07.
Open a file and the disassembly is tiny, the import table has five entries, and the strings are garbage? You are probably looking at a packed or protected binary: the real code is compressed/encrypted and only unpacks in memory at runtime. Before you can read it statically you need to recognize this — and tlk-hex gives you four quick signals in the Findings panel.
Entropy measures randomness on a 0–8 scale (bits per byte). Normal code sits around 6–6.8.
Compressed or encrypted data pushes toward 7.5–8, because packed bytes look like noise. tlk-hex shows an
entropy value per section in the Segments view and raises a finding when a non-resource section goes above
7.5.
| Entropy | Usually means |
|---|---|
| < 6.0 | Text, tables, uncompressed data |
| 6.0 – 7.0 | Ordinary machine code |
| > 7.5 | Compressed / encrypted — likely packed |
.rsrc alone is normal (it may hold PNGs or a zip). High
entropy in the code section is the red flag.Many packers leave their fingerprints in the section names. tlk-hex matches a built-in list and names the tool for you:
| Section | Packer |
|---|---|
UPX0 / UPX1 | UPX (free, easy to unpack) |
.aspack / .adata | ASPack |
.vmp0 / .vmp1 | VMProtect (virtualization) |
.themida / winlice | Themida / WinLicense |
.petite, .mpress, .fsg | Petite / MPRESS / FSG |
UPX is the friendly one: upx -d sample.exe unpacks it and you analyze the result normally.
VMProtect/Themida are a different league — they virtualize the code and are out of scope for static-only study.
A real GUI program imports dozens to hundreds of functions. If a non-DLL imports only a handful and you
also see LoadLibrary + GetProcAddress, the program is resolving its
real API at runtime to hide it from the import table. tlk-hex flags both the sparse import table and the
dynamic resolution pattern.
The imphash is an MD5 over the import table (each dll.function, in order). Two samples
built from the same source/toolkit usually share an imphash even if their bytes differ — so it's a cheap way to
cluster a malware family or spot siblings. tlk-hex computes it and shows it in Findings; copy it into a
threat-intel search to find related samples.
Even a packed file leaks a little. Before unpacking, skim:
VirtualAlloc + VirtualProtect
is the classic "allocate memory, unpack into it, make it executable" trio.jmp into freshly written memory (the "tail jump" to the original entry point / OEP).VirtualProtect(lpAddress, dwSize, flNewProtect, lpflOldProtect)) so the unpack stub reads clearly.Take any harmless program you own, pack a copy with UPX, and open both in tlk-hex:
upx -d, reopen, and verify the imports and strings come back.mov, call, jmp…rax, rcx…C3 = ret, etc.sub_X becomes its real name.