Back
Kevin Riedl

12 min read · 28 Sep 2026
Last reviewed

Next
Made on your device, with no Instagram connection. We copy the post link for Instagram’s Link sticker.

DeerFlow 2.0: Docker Setup, Sandboxes and Memory

DeerFlow 2.0 is an agent execution harness, not another model. The useful question is whether it can coordinate your tools, files and specialist agents with less custom integration work, while preserving boundaries you can test. The official ByteDance repository describes version 2 as a ground-up rewrite, rather than a small update to its earlier deep-research system.

Review scope: this guide reviews the public 2.x main documentation and implementation on . It is not launch-day coverage, a live installation test or a reproduced performance benchmark. Later revisions and older 2.0 checkouts can differ. Record your commit before applying the instructions.

What does DeerFlow actually orchestrate?

DeerFlow brings the agent runtime, tool execution, working files and reusable skills into one system built on LangGraph and LangChain. Its architecture documentation describes the agent middleware and file-based SKILL.md extensions. It does not replace LangGraph; it supplies a more assembled execution environment around those foundations.

Our assessment: it is worth evaluating when a task must research, inspect files, run code and return a verifiable deliverable across multiple steps. It is less compelling when a fixed script or ordinary queued workflow already solves the problem. For the category-level decision, use our agent harness engineering guide. This article focuses on operating DeerFlow itself.

A harness can reduce coordination code. Your team still defines what success means, which actions require approval, which data may enter a model, and who handles failures. Begin with one bounded workflow rather than giving a general-purpose assistant every company integration.

DeerFlow 2.0 Docker setup: the missing steps

A clone-and-start command is not a complete first-time installation. For the commands below, have Git, Make, Python 3, a suitable shell, a running Docker engine and Docker Compose available. The reviewed upstream instructions require Compose v2.24 or later. Use the configuration examples from the same checkout, not a snippet copied from an older release.

1. Clone the upstream repository and create configuration

set -eu
git clone https://github.com/bytedance/deer-flow.git
cd deer-flow
git rev-parse HEAD
docker compose version
make config

Keep the printed commit hash with your evaluation notes. The configuration bootstrap script creates config.yaml, .env and frontend/.env from their examples. It aborts when a root configuration already exists. That is not an instruction to delete a working configuration; back it up and review the repository’s upgrade path instead.

2. Configure a real model and choose the task sandbox

Edit the generated files before startup. Configure at least one supported model, use its actual provider credentials and review the frontend environment requirements. Keep secrets out of commits and shared diagnostic output. Do not leave placeholder model names or keys in an otherwise convincing installation example.

For an AIO container-backed evaluation, replace the existing sandbox: block in config.yaml with the following minimal selection. Do not append a second YAML key:

sandbox:
  use: deerflow.community.aio_sandbox:AioSandboxProvider

The sandbox configuration guide distinguishes the local provider from AIO container execution. Running the DeerFlow application in Docker does not, by itself, select an isolated task sandbox. This example targets local AIO containers, not an already configured remote provisioner. Review any existing provider settings before changing them.

3. Initialize the image and start the development stack

make docker-init
make docker-start

Open http://localhost:2026. The reviewed Makefile classifies make docker-start as Docker development startup. Use make docker-stop for that stack. The separate production path is make up, with make prod-logs and make down; a production command is not a security certification.

For development diagnostics:

make docker-logs

Before connecting a real workflow, run a synthetic smoke test: upload a tiny CSV with values 2 and 3, ask the agent to calculate the sum through its code tool, and require a saved result file containing 5. Independently inspect the file and tool trace. A chat answer saying “done” does not prove that code executed or that an artifact is available.

Troubleshoot startup before changing permissions

The upstream setup guide places config.yaml at the repository root and documents runtime data under .deer-flow, or a location set through DEER_FLOW_HOME. Check paths and mounted configuration before treating every failure as a model problem. The table below is a suggested diagnostic order, not a claim that these errors affect every installation.

DeerFlow Docker evaluation: first checks by symptom
SymptomCheck firstAvoid this shortcut
Configuration missing or rejectedRoot directory, actual mounted file, valid YAML and compatibility with the checked-out example.Deleting a working config or pasting duplicate top-level keys.
UI opens but the model call failsModel identifier, endpoint, credentials, provider access and gateway logs.Assuming UI availability proves the model configuration works.
Shell execution failsSelected sandbox provider, image availability and the provider’s container connectivity.Enabling unrestricted local shell execution merely to silence the error.
First code task appears stuckImage-pull progress, container startup and tool timeout evidence.Repeatedly submitting the same expensive task without checking logs.
Results disappear after restartConfigured state backends, persistent mounts, identities and artifact paths.Assuming every in-memory runtime component is durable.

Are DeerFlow subagents isolated from each other?

Separate conversation context does not imply a separate container. The reviewed subagent executor can carry the parent sandbox into delegated execution. Treat agents that share that sandbox as participants in a shared filesystem, even when their conversational histories are separate. A fresh model context is not a tenant-security boundary.

The current AIO provider implementation maps sandbox ownership using user and thread identity. It selects a remote backend when provisioner_url is configured; otherwise it uses the local container backend. Neither mechanism is a promise of one independent container per specialist subagent.

For a first workflow, give each parallel branch a distinct output filename, make source inputs read-only where supported, and let one synthesis step own the final report. Test whether one branch can overwrite another branch’s files. For separate customers, verify server-derived identity, authorization and storage boundaries rather than relying on names in a prompt.

The local provider is a different choice: its execution happens in the application’s environment rather than creating an additional task container. Where host-shell execution is disabled, do not remove that restriction casually. The OpenSandbox review covers sandbox infrastructure separately; a sandbox provider and a complete agent harness solve different layers.

How DeerFlow memory works, and why it still uses tokens

Separate three concerns: the current conversation, reusable memory across conversations, and runtime state needed to resume work. The shared memory configuration distinguishes enabling memory, injecting it into prompts and choosing a backend. “Memory is enabled” therefore does not prove that new facts are being extracted or that every unfinished run survives a restart.

In the reviewed configuration example, DeerMem’s max_injection_tokens belongs under memory.backend_config, with an example value of 2000. Its extraction model is separately configured; omitting the model settings leaves automatic extraction unavailable while non-LLM memory operations can still work. Check the configuration actually loaded by your revision.

Conversation compaction is another mechanism. The summarization middleware provides that layer; it is not the same as durable business memory. Summaries can lose details, and material added to a model’s context still consumes capacity. Do not translate “bounded injection” into “unlimited memory without token costs.”

Our suggested memory test has three parts. Store an innocuous preference, start a new conversation under the same identity, and verify retrieval. Then correct the preference and verify that the old value does not keep winning. Finally, repeat under another test identity and verify that the preference is unavailable there. Inspect logs and storage as well as the final answer.

Record the actual prompt and total provider usage. If memory increases cost without improving accepted outcomes, reduce what is injected or narrow the workflow. If the real problem is a memory layer for an existing agent stack, compare that narrower requirement with our OpenViking memory review rather than replacing the whole stack by default.

Custom DeerFlow skills: procedure, not permission

A useful first skill should encode a repeatable output contract, not an invitation to be generally autonomous. Using the file-based skill format described above, the following is an illustrative SKILL.md body for a source-linked vendor brief. Integrate and test it using your checkout’s skill discovery and approval rules; this is not a tested plugin package.

---
name: vendor-evidence-brief
description: Prepare a source-linked vendor brief for human review.
---

Use only the approved input documents and permitted sources.
Record a source URL or input filename for each factual claim.
Mark missing evidence as unknown; never invent a reference.
Write research branches to separate output files.
Produce a draft, not an approval or a published report.
Do not send messages, change accounts, or execute payments.

The distinction matters: “do not send messages” is an instruction, not a technical denial of messaging access. Enforce restricted actions through the tools and credentials available to the run. Treat third-party skill files, bundled scripts and dependencies as code or instructions requiring review, not as automatically trusted additions.

Keep reusable procedure in the skill and task-specific facts in the task inputs. That makes the skill easier to review, compare between revisions and test against both successful and malicious examples. Do not embed live customer secrets in reusable instructions.

A bounded workflow: vendor research to an evidence-backed draft

This is a proposed pilot design, not a reported DeerFlow customer deployment. Start with three approved vendor documents and a fixed question: which options meet a defined technical requirement? Restrict the deliverable to a draft comparison and an evidence file. Do not connect procurement, payment or outbound email actions.

A proposed acceptance contract for the first DeerFlow workflow
StageExpected artifactAcceptance check
Input preparationApproved files and explicit comparison criteria.Inputs contain no unnecessary secrets; scope and allowed sources are recorded.
Parallel researchA separate evidence file per vendor.Every factual row has an input reference or permitted source URL; unknowns stay unknown.
SynthesisOne draft comparison with limitations.Claims can be traced to the evidence files; missing evidence is not presented as failure or success.
Independent validationA structured validation result.Required files exist, machine-readable output parses, and sampled claims match their sources.
Human releaseAn approved or rejected draft.A named reviewer decides whether anything may be shared or acted upon.

Include a deliberately conflicting source, a missing document and an interrupted run in the acceptance set. Require the system to surface uncertainty rather than fabricate a clean answer. If a simple single-agent baseline performs better, keep it. Extra subagents are an implementation option, not an outcome metric.

DeerFlow production checklist: what must be demonstrated?

Older “no authentication” descriptions should not be treated as current. The reviewed authentication design includes first-run setup and authenticated user handling. Complete the initial administrator setup through the local /setup flow before making the service reachable by others, and test authorization with separate accounts.

The reviewed production Compose file defaults the published entry point to 127.0.0.1 and requires BETTER_AUTH_SECRET. Preserve loopback exposure during evaluation. Public access should follow an explicit deployment review, not a casual change to BIND_HOST.

Use the following as proposed release gates. They are not a statement that every control is supplied or enabled by DeerFlow.

Evidence to require before connecting production data or actions
ControlProof to collect
Identity and permissionsOne user cannot fetch another user’s threads, files or memory; privileged actions require an explicitly authorized identity.
Runtime containmentReview mounts, outbound network access, secrets and resource limits. Where AIO uses the host Docker socket, assess that privileged control path separately.
Persistence and recoveryRestart during a run, reconnect and inspect state. Verify backup restoration and that retries do not duplicate external actions.
Cost and cancellationProve spending limits, task deadlines, cancellation behavior and limits on parallel work against representative failure cases.
Skill and tool trustReview executable extensions, MCP launchers and skill changes. Untrusted content must not acquire administrative tool authority.
Upgrade ownershipPin application and image revisions, keep redacted traces, re-run acceptance tests and demonstrate rollback with a named operator.

For the final release review, use our software QA checklist before launch as a broader delivery checklist, then add the agent-specific cases above. Passing a happy-path demonstration is not enough evidence for account changes, payments or other irreversible actions.

Is DeerFlow worth adopting?

Measure cost per accepted task = (total model costs + total tool costs + total runtime costs + human review costs) / accepted tasks. Use consistent accounting units, count failed attempts within those totals, include memory extraction and delegated calls, and compare the same tasks against your existing baseline. Also record completion time, reviewer corrections, recovery rate and unsafe actions blocked. If no tasks are accepted, report that failure rather than a misleading average.

Choose a DeerFlow pilot when an integrated execution environment might remove real orchestration work. Keep a smaller framework or deterministic workflow when that meets the acceptance contract with less operational complexity. Our LangChain Deep Agents review helps frame the narrower framework option.

Wavect’s AI development service can help scope a bounded workflow, its permissions and acceptance tests. The TwinSoft AI case study is adjacent delivery context, not evidence of a DeerFlow deployment. To assess the fit for your system, discuss the workflow and its failure modes with Wavect.

DeerFlow 2.0 FAQ

What is DeerFlow 2.0?
DeerFlow 2.0 is ByteDance’s open-source agent harness, built with LangGraph and LangChain. It combines agent execution, subagents, tools, sandbox integration, memory and skills. It is not a new foundation model or a guarantee that a workflow will succeed.
Is git clone followed by make docker-start enough?
Not as a complete first-time setup. Prepare the configuration and environment files, configure a working model, choose the sandbox provider and initialize the sandbox image before starting the development stack.
Does every DeerFlow subagent get its own Docker container?
Do not assume that. Conversation isolation and sandbox isolation are different. Subagents can use the parent sandbox and shared files; verify the boundaries of the provider and deployment you actually run.
Why is DeerFlow memory not learning anything?
Check both the shared memory settings and the selected backend. In the reviewed DeerMem example, automatic extraction requires a configured extraction model. Also check credentials, identity scope, storage and logs. Existing memory injection and new memory extraction are separate operations.
Does DeerFlow memory remove token costs?
No. Injected memory, retrieved content, summaries and extraction calls can consume tokens. The relevant measurement is total cost per accepted task, including failed attempts and human review, not the size of one prompt.
Can DeerFlow be used in production?
Evaluate a pinned revision against your workload first. The repository has a separate production Docker command and authentication, but you still need tested permissions, persistence, containment, recovery, spending controls and operational ownership.

Build the product, not just the backlog

If this article maps to a real product decision, Wavect can help you scope, build, harden, or lead the software work with senior founder-level judgment.

Useful service paths:

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.

Back
Kevin Riedl

12 min read · 28 Sep 2026
Last reviewed

Next

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Free, double opt-in, no tracking pixels.