The interpreter is visible. The program has moved.
Traditional JavaScript analysis starts with functions, expressions, and branches. With VM-based obfuscation, much of the interesting program is represented as data. The JavaScript you can read may be the machinery that executes it, rather than the detection logic in a recognizable source form.
Our analysis of a supplied DataDome CAPTCHA-client snapshot found a custom JavaScript bytecode VM in Module 2. This article follows its architecture and the evidence recovered from it. The important question is not just whether the code can be made readable, but which claims the recovered behavior actually supports.
Scope: the findings concern this archived snapshot. The snapshot date is 9 September 2026. Analysis began on 8 September 2026. The release identifier is not established. We do not claim it represents the current production build, or that it is the exact implementation described in DataDome’s announcement.
What DataDome says publicly
DataDome’s VM-based obfuscation announcement describes compiling detection logic into custom bytecode and executing it with a browser-side interpreter. It identifies a program counter, stack, internal memory, and a proprietary instruction set, alongside regeneration of the instruction mapping and implementation. The announcement names Device Check and Slider and positions VM obfuscation alongside existing dynamic and WASM protections.
Those are vendor descriptions. Our supplied bundle establishes a CAPTCHA-client analysis context. It does not establish a release-to-release link with that announcement. The useful comparison is architectural, rather than an assertion that every detail or every marketing claim has been independently verified.
Methodology
The underlying investigation separated module extraction, bytecode decoding, control-flow recovery, semantic reconstruction, and runtime validation. For this publication, we reviewed the frozen Module 2 reports, machine-readable disassembly, and stored validation records. We did not execute the original bundle or repeat its browser experiments.
That distinction matters. A reconstruction report can explain behavior. A trace can establish what ran under one environment. Neither alone proves all possible paths or the server’s decision process.
We publish a sanitized structural summary with the analyzed module’s SHA-256 digest and aggregate record counts. It contains no vendor implementation, addresses, handler mappings, runtime patches, or encoded payloads.
| Evidence inspected | Recorded result | What it establishes |
|---|---|---|
| Decoded payload metadata | 230,027 bytes | The size reported by the archived decoder for this snapshot |
| Function metadata | 184 function entries | Recovered function boundaries, not 184 independently validated behaviors |
| Decoder problem list | 0 entries | No recorded decoding problems. This does not prove complete semantic correctness |
| Partial trace validation | 251 stack-entry checks, 0 mismatch entries | Agreement at those recorded checkpoints, not exhaustive execution coverage |
The structural counts above were read from machine-readable artifacts. Later semantic reports supply additional interpretation. Those interpretations remain bounded by the archive’s explicit unknowns.
Findings
1. A custom VM and WebAssembly are distinct layers
Module 2’s custom interpreter is implemented in JavaScript and executes its own bytecode. The archive separately documents an embedded WebAssembly component in Module 1. Treating every opaque binary-looking object as WASM would collapse two different execution models into one.
For the custom VM, the analytical unit is an instruction and its effect on VM state. For the separate WASM component, the relevant execution model is WebAssembly. Their presence in the same bundle does not make them interchangeable.
This is a conceptual diagram of the recovered stages, not a vendor opcode specification.
2. Recovering the language is different from recovering the program
An opcode handler describes how an instruction changes state. It does not by itself explain the higher-level purpose of every program that uses that instruction.
The archive’s reconstruction proceeds from low-level behavior to control flow and then to selected semantic helpers. Its decompiler history preserves unresolved structures rather than silently replacing them with plausible-looking JavaScript. That is an important research discipline: readable output is useful only if the transformation preserves the behavior being studied.
Two sources of control flow also need separating. The interpreter dispatches bytecode instructions. The protected program can contain its own flattened dispatcher. The final reports describe a recovered 43-state program dispatcher. That count is specific to the supplied sample, not a universal property of DataDome’s VM.
3. Values need lifetimes, not just names
The archive records a substantial correction around frame and operand-stack semantics. If a value is pushed and its original storage slot is later overwritten, a reconstruction must preserve the earlier value. Replacing every read with a convenient name for the current slot can make the reconstructed program mean something different.
Some scratch locations are reused for different observations on different routes. Naming a write site is defensible when its producer is understood. Assigning one permanent semantic meaning to the entire slot can be misleading.
This is where an analysis becomes more than a beautified source dump. It needs a model of evaluation order, state, and dataflow, with checks against execution evidence.
4. Browser observations become mixed state
The reconstructed operations span browser identity, layout and rendering, canvas/WebGL, timing, runtime inspection, and emulation-related environment observations. Those categories describe the APIs and behavior visible in the reconstruction. They do not disclose a server-side rulebook.
The reports describe observations entering an intermediate mixing layer and a packed accumulator before an encoded result is returned to the surrounding bundle. At this layer, there is no simple, fully attributed object of named detection fields.
Recognizing an API access establishes that the program uses it. Establishing its contribution to the final result requires tracing its consumers. Establishing its contribution to a server decision is a further step.
5. A changed observation need not change the route
The recorded comparison runs followed the same observed program-dispatch route while a runtime-inspection observation and its downstream table effect differed. This is useful because it separates data changes from control-flow changes.
The archive does not establish which final packed byte, if any, directly represents that difference. It also records timing and random variation affecting independent outputs. Comparing two whole output strings is therefore insufficient to attribute every difference to the one observation of interest.
The defensible conclusion is narrower: a recorded observation reached internal state. It is not proof of an exact output field, a detection weight, an allow/block decision, or a working bypass.
What VM obfuscation changes for analysis
VM obfuscation adds a language-recovery problem before ordinary program analysis can begin. The researcher has to distinguish interpreter machinery, protected-program control flow, data transformation, and browser interactions. Decoding instructions is an important milestone, but it does not finish semantic attribution.
Our snapshot illustrates why structural recovery, behavioral validation, and causal explanation are separate deliverables. The first identifies the execution model. The second checks selected behavior. The third needs evidence connecting a specific cause to a specific result.
None of this measures the vendor’s regeneration cadence, the time required by another researcher, or the effectiveness of the protection across deployments.
Limitations
- The publication covers the 9 September 2026 snapshot. Its release identifier is not confirmed. Current deployments may differ.
- The archive contains stored runtime evidence, but its experiments were not repeated for this publication. Partial traces are not full path coverage.
- Explanatory names are reconstruction labels, not recovered original source names.
- Some mixing operations and individual packed-output meanings remain unattributed in the reports.
- Whole-output differences can include timing and random variation. They are not a one-signal causal map.
- Client-side behavior does not reveal all server-side features, weights, context, or decision rules.
- The module digest identifies the analyzed artifact but cannot make unpublished input and tooling independently reproducible.
Reproduction
The structural summary exposes the publication’s aggregate evidence and boundaries. Reproducing the vendor-specific counts requires the same private snapshot and analysis artifacts. This article does not claim a complete public reproduction package.
For a safe educational exercise, use an interpreter and program you own: record instruction boundaries, stack depths, frame writes, and observable outputs. Check that your lifted representation preserves values across overwrites and that only equivalent control-flow regions are simplified. A successful toy exercise demonstrates the method, not the correctness of the DataDome reconstruction.
No opcode mapping, bytecode decoder, exact probe implementation, patch, payload builder, or CAPTCHA-solving workflow is published here. The investigation is about understanding execution and stating what the evidence can establish.
Sources and disclosure
The vendor framing comes from DataDome’s public changelog. Our sample-specific findings come from the owner’s supplied research archive and its frozen Module 2 artifacts, with aggregate evidence linked above. This is independent research and is not affiliated with or endorsed by DataDome.
Found a mistake or a result you cannot reproduce?
Send a correctionReference: /articles/datadome-vm-obfuscation/