reverse engineering from scratch

Learn to read
machine code.

Even if you've never seen assembly, you can start here. From concepts to solving a real function line by line — trying every step inside tlk-hex.

module 01

Fundamentals beginner

You write a program, the compiler turns it into machine code: the bytes the processor runs directly. The source code is usually not available — all we have is those bytes. Reverse engineering is working backwards from those bytes to understand what the program does.

What does a disassembler do?

The processor reads bytes and runs them as instructions. For example the bytes 48 83 EC 28 actually mean "reserve 0x28 bytes on the stack". A disassembler turns those bytes into human-readable assembly:

bytes → assembly
48 83 EC 28      sub     rsp, 28h    ; reserve 40 bytes on the stack

tlk-hex decodes the whole file this way, then ties the functions, calls and data together.

A few basic concepts

  • Register: tiny, very fast variables inside the processor. On x64 they're named rax, rbx, rcx…
  • Memory address: every byte has a number, written in hexadecimal like 0x140001008.
  • Instruction: one step for the processor — mov, add, call…
  • Function: a sequence of instructions that work together; called from somewhere with call, returns with ret.
Why hexadecimal? Bytes range 0–255 and the range 00–FF fits exactly into two hex digits. That's why addresses and values are always written in hex. In tlk-hex, H converts a number to decimal.
module 02

File structure beginner

An .exe isn't just code. For Windows to load it into memory correctly it has a layout: the PE (Portable Executable) format. On Linux the equivalent is ELF.

Sections

The file is split into sections by purpose. In tlk-hex, Shift+F7 shows them all:

SectionWhat's inside
.textExecutable code. Functions live here.
.rdataRead-only data: strings, import table, constants.
.dataWritable global variables.
.rsrcResources: icons, version info, dialogs.
.pdatax64 exception/unwind info — helps find function boundaries.

Key points

  • Entry point: the first instruction run when the program starts. Ctrl+E in tlk-hex.
  • Imports: functions the program borrows from other DLLs (CreateFileW, send…). The strongest clue to what it does.
  • Exports: functions a DLL exposes to the outside.
  • Image base: the start address where the file loads in memory (0x140000000 for most x64 EXEs).
RVA vs offset: the position in the file (file offset) differs from the address in memory (RVA + image base). tlk-hex shows both and does the conversion for you — you see both in the status bar.
module 03

Reading assembly intermediate

It looks scary but a few dozen instructions cover 90% of the work. First, the registers:

Registers

64-bit32-bitTypical use
raxeaxReturn value; general math
rcx rdx r8 r9ecx edxFirst 4 arguments to a function (x64)
rbx rsi rdiebx esi ediPreserved general variables
rspespStack pointer
rbpebpStack frame base
ripeipAddress of the next instruction

Smaller parts of a register are also used: rax → eax (32) → ax (16) → al (8 bit).

The most common instructions

InstructionMeaning
mov a, ba = b — copy a value
lea a, [b]a = address of b — compute an address (doesn't always read memory)
add / subAdd / subtract
xor a, aZero a register (XOR with itself = 0)
cmp a, bCompare (a - b) and set flags
test a, aCheck whether a is zero (very common)
jmpUnconditional jump
jz / jeJump if equal / zero  ·  jnz/jne: if not
callCall a function
retReturn from a function
push / popPush onto / pop off the stack

Calling convention (x64 Windows)

When a function is called, arguments go in order into rcx, rdx, r8, r9; extras go on the stack. The return value is in rax. So when you see this:

.text — a call
mov     rcx, 1
lea     rdx, aHello         ; "hello"
call    WriteMessage       ; WriteMessage(1, "hello")
mov     ebx, eax           ; save the return value

…you're actually reading the call WriteMessage(1, "hello").

Stack variables

Local variables live on the stack. tlk-hex names them automatically as var_18 (local) and arg_8 (argument); the text [rsp+88h+var_18] means "a local variable of this function". Hover over it and press N to give it a meaningful name.

module 04

First analysis intermediate

Now let's read a real function. Take the example from the home page and solve it line by line:

IDA View-A · ParseConfigFile
140001008 ParseConfigFile proc near
140001008   mov     r11, rsp
14000100B   sub     rsp, 88h              ; reserve stack space
140001012   mov     rax, cs:__security_cookie  ; stack protection
140001019   xor     rax, rsp
14000101C   mov     [rsp+88h+var_18], rax
140001053   lea     rcx, aConfigLoaded    ; "config loaded: %s"
14000105A   mov     edx, esi
140001073   call    LogMessage            ; LogMessage("config loaded: %s", esi)
140001080   call    cs:__security_check_cookie
140001085   add     rsp, 88h              ; give the space back
14000108C   retn

What did we read?

  • The prologue (sub rsp + security_cookie): the standard opening of every compiler-generated function. Protection against stack overflows. The logic isn't here.
  • The middle: lea rcx, aConfigLoaded takes the address of a string, call LogMessage prints it. So this function logs something.
  • The epilogue (security_check + add rsp + retn): the standard closing.

So once you strip the noise, the essence of the function is one line: log the "config loaded" message. Half of real analysis is recognising these standard patterns and focusing on what matters.

tlk-hex tip: wondering where a string is used? Hover over it and press X — every place that uses it is listed. The reverse works too: double-click a call target to step into that function.

Where to start?

  • Strings (Shift+F12): text like "invalid password", "connecting to..." gives away the program's intent. Press X on an interesting string to reach the code using it.
  • Imports: see socket/connect and there's networking; see CryptEncrypt and there's encryption.
  • Entry point (Ctrl+E): follow the flow from the very top.
  • Findings tab: tlk-hex already summarises suspicious/interesting spots — a quick map.
module 05

Recover the flow intermediate

Press Space for the graph. The arrows between blocks are the program's logic. Learn a few patterns:

if / else

A cmp + conditional jump is a branch. Two arrows leave the block: green = condition true, red = false.

a decision
cmp     eax, 0
jz      loc_error      ; if eax == 0 jump to error
; ... falling through here means eax != 0 (success path)

In C: if (eax == 0) goto error;

loop

When you see an arrow in the graph that goes back up, that's a loop: the block jumps back to an earlier one. It usually ends with a counter (inc/dec) and a cmp.

switch

Branching many ways based on a value. tlk-hex resolves jump tables automatically and adds comments like jumptable ... case 3 — you read directly which value goes where.

Shortcut: instead of reading block by block, press F5. tlk-hex converts the flow into a C-like summary with if/goto for you. Not exact, but it conveys the logic fast.
module 06

Practice advanced

Reading doesn't teach; solving does. Safe and legitimate ways to begin:

  • Compile your own program, then analyse it. Write a small C/C++ program, open it in tlk-hex. Since you know the source, you see the assembly mapping one to one — the fastest way to learn.
  • CTF (Capture The Flag): the "reversing" categories of security competitions are designed exactly for this; they're files made for you to solve.
  • Crackmes: practice programs the author deliberately shares as "solve me". The goal here is to understand a mechanism — not to crack someone's commercial software.
  • Malware analysis: advanced. Examine samples in an isolated environment, statically only (tlk-hex never runs them, so reading is safe).
Where's the line? Reverse engineering is legitimate for learning, inspecting your own code, security research and authorized testing. Removing the licence/copy protection of software you didn't buy (cracking) or infringing someone else's rights is outside that. Work on files you're authorized for.

A working routine

  • First get an overview: strings, imports, findings. What does the program roughly do?
  • Enter from an interesting clue (X from a string/import into code).
  • Read the function in the graph, strip the standard patterns; verify with F5.
  • Name things (N) and leave notes (;). As you understand, the file becomes readable.
  • Save with Ctrl+W; continue tomorrow where you left off.
module 07

Solving a full EXE, end to end intermediate

So far we've seen the pieces separately: strings, imports, xrefs, assembly, graph and pseudocode. Now we'll combine them into one real analysis workflow. Don't rush — the goal of this module isn't speed, it's learning how to ask a binary questions.

We'll pretend we have no source code and answer these only from the file itself:

  • What input does the program care about?
  • Which function makes the decision?
  • Which condition leads to which outcome?
  • How are the functions wired together?
  • And most importantly: how do we prove what we found is correct?
For this exercise we use a tiny EXE we compile ourselves. That keeps the analysis fully safe, legal and verifiable: at the end we can check our findings against the real source.

1 · Build the target program

First we write the program we'll analyse. Create a file sample.c:

sample.c
#include <stdio.h>
#include <string.h>

static int classify(const char *mode)
{
    if (strcmp(mode, "debug") == 0)
        return 2;
    if (strcmp(mode, "safe") == 0)
        return 1;
    return 0;
}

int main(int argc, char **argv)
{
    const char *mode = argc > 1 ? argv[1] : "safe";
    int result = classify(mode);
    printf("mode=%s result=%d\n", mode, result);
    return result;
}

For now, forget what it does. We only use the source to create our target; once the EXE exists we look at it purely as a binary. Compile it (whichever you have):

MSVC (Developer Command Prompt)
cl /O2 /Fe:sample.exe sample.c
MinGW-w64
gcc -O2 -o sample.exe sample.c

We now have sample.exe. From here on, close the source.

2 · Open it and map it first

Drag sample.exe onto tlk-hex. Auto-analysis starts and recovers sections, functions, calls, strings, xrefs and imports.

When it finishes, don't dive into assembly yet. Get a feel for the file first. In reverse engineering the first goal isn't detail, it's context — reading instructions one by one without knowing where you are just wears you out.

3 · Look at strings first

Open Strings with Shift+F12. You'll look for text like:

Strings
debug
safe
mode=%s result=%d

These three lines already tell us a lot before reading a single instruction. debug and safe are probably values compared against user input; mode=%s result=%d suggests the program prints a result at the end.

But these are still only hypotheses. The golden rule: don't treat something as true just because it looks reasonable; verify it with xrefs and control flow.

4 · Go from string to code (xref)

Select the debug string and press X. This shows where it's used (cross-references). Jump to a code reference. Do the same for safe.

You'll see both strings' references land in the same function or very close code. A strong clue:

This function most likely compares the incoming value against "debug" and "safe".

5 · Check imports, combine the evidence

Now look at the functions the program borrows from outside. Names vary a little by compiler, but you'll see:

ImportWhat it does
strcmpCompares two C strings (returns 0 if equal)
printfPrints formatted text

Now three independent pieces of evidence have piled up: "debug", "safe" and strcmp. Together they point to a strong likelihood: the program compares a string against "debug" and "safe". We still won't be sure until we see it in assembly.

6 · Read the deciding function

Go to the start of the function you reached from the debug xref. Its auto name might be sub_140001000 — we don't know its real name yet. Inside you'll see a pattern like this (exact instructions vary by compiler; what matters is reading the pattern):

comparison pattern
lea     rdx, aDebug        ; 2nd arg: "debug"
mov     rcx, rbx          ; 1st arg: the input
call    strcmp            ; strcmp(input, "debug")
test    eax, eax          ; is the result 0? (equal?)
jz      loc_debug         ; if equal, jump to the debug block

The first two lines prepare strcmp's arguments. strcmp returns 0 when the strings are equal; test eax, eax checks whether the result is zero and jz (jump if zero) goes to that block. So the logic is clear:

C equivalent
if (strcmp(input, "debug") == 0)

7 · Give the function a meaningful name

We roughly know what it does now. On the function, press N and rename sub_140001000 to classify_mode.

Naming isn't a luxury, it's the heart of the method. At first you see dozens of sub_140001000; as you turn them into classify_mode, print_result, load_config and so on, the file gets more readable every pass.

8 · Switch to the graph, see the flow

Inside the function, press Space; instead of text you get the control-flow graph. For this example we roughly expect this decision tree:

expected decision structure
                input
                  │
                  ▼
          input == "debug"?
           ┌──────┴──────┐
          yes           no
           │             │
           ▼             ▼
       return 2    input == "safe"?
                    ┌─────┴─────┐
                  yes          no
                    │           │
                    ▼           ▼
                return 1    return 0

What looks tangled in assembly reads at a glance in the graph. Follow the no (false) path of the first comparison and after a while you'll find a second comparison using the "safe" string — same logic:

second comparison
lea     rdx, aSafe
mov     rcx, rbx
call    strcmp
test    eax, eax
jz      loc_safe          ; if (strcmp(input, "safe") == 0)

9 · Follow the return values

Now check which value each path returns. On x64 Windows the integer return value comes back in eax:

return values
mov     eax, 2
ret                     ; return 2;  (debug path)

mov     eax, 1
ret                     ; return 1;  (safe path)

xor     eax, eax
ret                     ; return 0;  (xor = zero it; default path)

xor eax, eax zeroes the register — i.e. return 0;. Recognising this little pattern pays off; you'll see it everywhere.

10 · Reconstruct the logic and compare with F5

Combining the evidence, the binary is telling us:

the logic we recovered
int classify_mode(const char *input)
{
    if (strcmp(input, "debug") == 0)
        return 2;
    if (strcmp(input, "safe") == 0)
        return 1;
    return 0;
}

We recovered the function's behaviour without the source — and that's exactly the point of reverse engineering: not getting the source back verbatim, but understanding the behaviour well enough.

Now press F5; tlk-hex turns the same function into C-like pseudocode and shows something very close to what you worked out by hand. It's fast, but a warning:

Decompiler output is not the real source code. Compiler optimizations, lost type information and register reuse can make it wrong or incomplete. Always verify important conclusions in assembly and the graph. Ideal order: get the idea from pseudocode → check the flow in the graph → prove it in assembly.

11 · Who calls it + the output chain

On classify_mode, press X again — this time you see who calls the function. You'll probably land in main or near it:

call site
mov     rcx, rbx          ; input
call    classify_mode
mov     esi, eax          ; save the return value (used later)

Then back to Strings, follow mode=%s result=%d with X; nearby you'll see a printf call. So we've recovered the whole chain:

program flow
command-line input  →  classify_mode  →  0 / 1 / 2  →  printf  →  exit

12 · Name, comment, save

Make your findings permanent:

  • N to name: sub_140001000 → classify_mode, aDebug → mode_debug, aSafe → mode_safe…
  • ; to comment key lines: "compares the incoming mode with 'debug'" or "default path returns 0".
  • Ctrl+W to save. tlk-hex keeps names/comments/work separately; the original EXE is untouched and reopening resumes where you left off.

Reverse engineering isn't only reading a binary; it's recording what you find in an orderly way. A well-named database looks completely different — far more readable — a few hours in.

13 · Write the result in a few sentences

A good session ends with a short conclusion. For this example:

analysis result
sample.exe takes a "mode" value from the command line.
  mode == "debug"  →  result = 2
  mode == "safe"   →  result = 1
  any other value  →  result = 0
The result is printed with printf as "mode=%s result=%d".

We recovered the program's core behaviour without ever looking at the source. That was the whole point.

The workflow you'll reuse on every binary

This loop works on almost any file — no need to memorize it, it becomes second nature after a few runs:

methodology
 1. Open the file, wait for auto-analysis
 2. Look at sections / imports / strings (context)
 3. Pick an interesting string or import
 4. Go to code with xref (X)
 5. Study the function's graph (Space)
 6. Follow the call targets
 7. Extract the conditions and return values
 8. Quick check with F5 pseudocode
 9. Verify in assembly
10. Name the functions (N)
11. Add comments (;)
12. Save (Ctrl+W)
13. Write the result in a few sentences

What to look at first in an unknown EXE

Strings are the strongest clues. If you see these, stop and follow string → X → function:

notable strings
error   failed   password   token   config
login   connect  http       registry file   debug

Imports reveal which capabilities the program uses:

Import groupWhat it points to
CreateFileW · ReadFile · WriteFileFile operations
RegOpenKeyExW · RegSetValueExWRegistry usage
connect · send · recvNetworking
CryptEncrypt · CryptDecryptEncryption
A single import proves nothing — weigh the evidence together. If you don't understand a function, ask two questions: who does it call? and who calls it? Xrefs answer both.

The limit of static analysis

tlk-hex currently does mostly static analysis — it inspects the program without running it. So assembly analysis, function discovery, xrefs, string/import analysis, the graph, pseudocode, PDB symbols, hex inspection and patch preparation all work comfortably. But some things can't always be seen with pure static analysis:

cases that need dynamic analysis
values computed at runtime
encrypted code decrypted in memory
self-modifying code
runtime unpacking (packers that unpack while running)

In those cases dynamic tools like a debugger come in. Still, most of the work starts static; you use static analysis to find what to look for, then switch to dynamic if needed.

Final task — try it yourself

Recompile the sample with a small change (try to analyse it without looking at the source again):

sample.c — add a line
if (strcmp(mode, "admin") == 0)
    return 3;

Open the new EXE in tlk-hex and find, from the binary only:

  • Which is the new string?
  • Which function uses it?
  • After which condition does it run?
  • Which value is returned?
  • Where is the new branch in the graph?
  • Does the pseudocode catch the change correctly?

If you can find all of these from the binary, you've completed your first real reverse-engineering workflow.

The short rule: guess → find evidence → name → verify → take notes.
When a binary first opens it looks like one huge file of meaningless addresses. Every string, xref and function you solve turns a small part of that unknown into something meaningful. The goal isn't to memorize all the assembly; it's to rebuild the program's behaviour step by step.
module 08

Patching: change what a program does intermediate

Sometimes you don't want to find the right password/serial — you want to change the program's behaviour: e.g. flip a check so it "always passes". That's patching: editing the bytes on disk. tlk-hex does this safely — it never touches the original file and keeps changes separately.

This skill is for your own software, CTFs, crackmes and authorized analysis — not for removing the protection of software you didn't buy.

When to patch

  • When satisfying the check (a keygen) is hard, but changing a single branch is easy.
  • When you want to quickly disable a check and study what's below it.
  • When you want to temporarily change a value/flow to test a hypothesis.

Find the conditional jump

Most checks end in a cmp/test + a conditional jump. A typical "deny on failure" pattern:

a decision
cmp     eax, 0DEADBEEFh
jnz     loc_denied        ; if NOT equal, jump to deny
; ... falling through here is the "granted" path

There are a few ways to change the flow — each is a one-byte edit:

GoalHowByte
Invert the conditionjnz → jz (or vice versa)75 ↔ 74
Never take the jumpNOP the conditional jump90 90
Always jumpMake it unconditional: jnz → jmp75 → EB
74 = jz/je, 75 = jnz/jne, 90 = nop, EB = short jmp. Memorizing these four bytes speeds you up a lot.

Patch it in Hex View

  1. Place the cursor on the instruction to change (e.g. jnz loc_denied).
  2. Switch to the Hex View tab; the cursor lands on that byte. Press F2 to enter edit mode.
  3. Type the byte: change 75 to 74 (or whatever you need). F2 / Esc to leave.
  4. Back in IDA View the instruction now reads jz — the logic is inverted.

Apply to disk (original untouched)

The patch lives in the database for now. To produce a permanent file:

  • File → Apply patches to file… — writes a new exe with the patched bytes (original stays separate).
  • File → Create DIF file… — exports the changed bytes as text (to share / document).
  • File → List patched bytes — shows everything you changed in one list.
Patching changes behaviour: keep a backup first, run and verify the result, and only do it on files you're authorized for.

Exercise

Two of the practice crackmes are made for this module:

  • crackme05 — the right key is hard to find (target 0DEADBEEFh). Find the branch after the cmp and patch it → "Access granted".
  • crackme06 — the password is a stack string (not in Strings); recover it by reading the byte stores.

Grab both from the crackmes release. The other levels fit the workflow in Module 07.

module 09

Packers & triage: when the code is hidden intermediate

Open a file and the disassembly is tiny, the import table has five entries, and the strings are garbage? You are probably looking at a packed or protected binary: the real code is compressed/encrypted and only unpacks in memory at runtime. Before you can read it statically you need to recognize this — and tlk-hex gives you four quick signals in the Findings panel.

1. Entropy — how random is a section?

Entropy measures randomness on a 0–8 scale (bits per byte). Normal code sits around 6–6.8. Compressed or encrypted data pushes toward 7.5–8, because packed bytes look like noise. tlk-hex shows an entropy value per section in the Segments view and raises a finding when a non-resource section goes above 7.5.

EntropyUsually means
< 6.0Text, tables, uncompressed data
6.0 – 7.0Ordinary machine code
> 7.5Compressed / encrypted — likely packed
A high-entropy .rsrc alone is normal (it may hold PNGs or a zip). High entropy in the code section is the red flag.

2. Packer signatures — the section names give it away

Many packers leave their fingerprints in the section names. tlk-hex matches a built-in list and names the tool for you:

SectionPacker
UPX0 / UPX1UPX (free, easy to unpack)
.aspack / .adataASPack
.vmp0 / .vmp1VMProtect (virtualization)
.themida / winliceThemida / WinLicense
.petite, .mpress, .fsgPetite / MPRESS / FSG

UPX is the friendly one: upx -d sample.exe unpacks it and you analyze the result normally. VMProtect/Themida are a different league — they virtualize the code and are out of scope for static-only study.

3. Sparse imports — the capability is missing

A real GUI program imports dozens to hundreds of functions. If a non-DLL imports only a handful and you also see LoadLibrary + GetProcAddress, the program is resolving its real API at runtime to hide it from the import table. tlk-hex flags both the sparse import table and the dynamic resolution pattern.

4. imphash — fingerprint the import table

The imphash is an MD5 over the import table (each dll.function, in order). Two samples built from the same source/toolkit usually share an imphash even if their bytes differ — so it's a cheap way to cluster a malware family or spot siblings. tlk-hex computes it and shows it in Findings; copy it into a threat-intel search to find related samples.

Reading what's left

Even a packed file leaks a little. Before unpacking, skim:

  • The few imports that are present — VirtualAlloc + VirtualProtect is the classic "allocate memory, unpack into it, make it executable" trio.
  • The entry point — a packer stub is short and ends with a jmp into freshly written memory (the "tail jump" to the original entry point / OEP).
  • API call comments — tlk-hex annotates known calls with their parameters (e.g. VirtualProtect(lpAddress, dwSize, flNewProtect, lpflOldProtect)) so the unpack stub reads clearly.

Exercise

Take any harmless program you own, pack a copy with UPX, and open both in tlk-hex:

  1. Compare the Segments entropy — watch the code section jump above 7.5 in the packed copy.
  2. Confirm the UPX packer finding and the sparse import table.
  3. Run upx -d, reopen, and verify the imports and strings come back.
Only pack/unpack software you own or are authorized to analyze. Never try to defeat the protection of software you didn't buy.
reference

Glossary

Disassembler
A tool that turns machine-code bytes into readable assembly. tlk-hex is a disassembler.
Decompiler
A tool that tries to turn assembly into higher-level (C-like) code. tlk-hex's F5 is like a simple decompiler.
Assembly
The human-readable form of processor instructions — mov, call, jmp…
Register
Very fast small storage inside the processor: rax, rcx…
Opcode
The raw byte(s) representing an instruction. C3 = ret, etc.
RVA
Relative Virtual Address — a memory address relative to the image base.
Entry point
The address of the first instruction run when the program starts.
Import
An external function the program uses from another DLL.
Xref (cross-reference)
References to an address/name — "where is this called/used from?"
Stack
The LIFO memory area functions use for locals and return addresses.
Calling convention
The rule defining which register/stack order arguments are passed in.
Prologue / Epilogue
A function's standard opening/closing code (stack setup, protection).
Basic block
A straight-line instruction sequence with no branching; one box in the graph.
PDB / Symbol
A helper file containing function names. Once loaded, sub_X becomes its real name.
Patch
Changing bytes in the file. In tlk-hex via F2; without breaking the original.
Packer
Protection that compresses/encrypts code and unpacks at runtime. Shows up as high entropy.
Entropy
A 0–8 measure of randomness. > 7.5 in a code section usually means packed/encrypted.
imphash
An MD5 fingerprint of the import table; shared across samples from the same toolkit/family.