In this piece
DeerFlow 2.0: Docker Setup, Sandboxes and Memory
DeerFlow 2.0 is an agent execution harness, not another model. The useful question is whether it can coordinate your tools, files and specialist agents with less custom integration work, while preserving boundaries you can test. The official ByteDance repository describes version 2 as a ground-up rewrite, rather than a small update to its earlier deep-research system.
Review scope: this guide reviews the public 2.x main documentation and implementation on . It is not launch-day coverage, a live installation test or a reproduced performance benchmark. Later revisions and older 2.0 checkouts can differ. Record your commit before applying the instructions.
What does DeerFlow actually orchestrate?
DeerFlow brings the agent runtime, tool execution, working files and reusable skills into one system built on LangGraph and LangChain. Its architecture documentation describes the agent middleware and file-based SKILL.md extensions. It does not replace LangGraph; it supplies a more assembled execution environment around those foundations.
Our assessment: it is worth evaluating when a task must research, inspect files, run code and return a verifiable deliverable across multiple steps. It is less compelling when a fixed script or ordinary queued workflow already solves the problem. For the category-level decision, use our agent harness engineering guide. This article focuses on operating DeerFlow itself.
A harness can reduce coordination code. Your team still defines what success means, which actions require approval, which data may enter a model, and who handles failures. Begin with one bounded workflow rather than giving a general-purpose assistant every company integration.
DeerFlow 2.0 Docker setup: the missing steps
A clone-and-start command is not a complete first-time installation. For the commands below, have Git, Make, Python 3, a suitable shell, a running Docker engine and Docker Compose available. The reviewed upstream instructions require Compose v2.24 or later. Use the configuration examples from the same checkout, not a snippet copied from an older release.
1. Clone the upstream repository and create configuration
set -eu
git clone https://github.com/bytedance/deer-flow.git
cd deer-flow
git rev-parse HEAD
docker compose version
make configKeep the printed commit hash with your evaluation notes. The configuration bootstrap script creates config.yaml, .env and frontend/.env from their examples. It aborts when a root configuration already exists. That is not an instruction to delete a working configuration; back it up and review the repository’s upgrade path instead.
2. Configure a real model and choose the task sandbox
Edit the generated files before startup. Configure at least one supported model, use its actual provider credentials and review the frontend environment requirements. Keep secrets out of commits and shared diagnostic output. Do not leave placeholder model names or keys in an otherwise convincing installation example.
For an AIO container-backed evaluation, replace the existing sandbox: block in config.yaml with the following minimal selection. Do not append a second YAML key:
sandbox:
use: deerflow.community.aio_sandbox:AioSandboxProviderThe sandbox configuration guide distinguishes the local provider from AIO container execution. Running the DeerFlow application in Docker does not, by itself, select an isolated task sandbox. This example targets local AIO containers, not an already configured remote provisioner. Review any existing provider settings before changing them.
3. Initialize the image and start the development stack
make docker-init
make docker-startOpen http://localhost:2026. The reviewed Makefile classifies make docker-start as Docker development startup. Use make docker-stop for that stack. The separate production path is make up, with make prod-logs and make down; a production command is not a security certification.
For development diagnostics:
make docker-logsBefore connecting a real workflow, run a synthetic smoke test: upload a tiny CSV with values 2 and 3, ask the agent to calculate the sum through its code tool, and require a saved result file containing 5. Independently inspect the file and tool trace. A chat answer saying “done” does not prove that code executed or that an artifact is available.
Troubleshoot startup before changing permissions
The upstream setup guide places config.yaml at the repository root and documents runtime data under .deer-flow, or a location set through DEER_FLOW_HOME. Check paths and mounted configuration before treating every failure as a model problem. The table below is a suggested diagnostic order, not a claim that these errors affect every installation.
| Symptom | Check first | Avoid this shortcut |
|---|---|---|
| Configuration missing or rejected | Root directory, actual mounted file, valid YAML and compatibility with the checked-out example. | Deleting a working config or pasting duplicate top-level keys. |
| UI opens but the model call fails | Model identifier, endpoint, credentials, provider access and gateway logs. | Assuming UI availability proves the model configuration works. |
| Shell execution fails | Selected sandbox provider, image availability and the provider’s container connectivity. | Enabling unrestricted local shell execution merely to silence the error. |
| First code task appears stuck | Image-pull progress, container startup and tool timeout evidence. | Repeatedly submitting the same expensive task without checking logs. |
| Results disappear after restart | Configured state backends, persistent mounts, identities and artifact paths. | Assuming every in-memory runtime component is durable. |
Are DeerFlow subagents isolated from each other?
Separate conversation context does not imply a separate container. The reviewed subagent executor can carry the parent sandbox into delegated execution. Treat agents that share that sandbox as participants in a shared filesystem, even when their conversational histories are separate. A fresh model context is not a tenant-security boundary.
The current AIO provider implementation maps sandbox ownership using user and thread identity. It selects a remote backend when provisioner_url is configured; otherwise it uses the local container backend. Neither mechanism is a promise of one independent container per specialist subagent.
For a first workflow, give each parallel branch a distinct output filename, make source inputs read-only where supported, and let one synthesis step own the final report. Test whether one branch can overwrite another branch’s files. For separate customers, verify server-derived identity, authorization and storage boundaries rather than relying on names in a prompt.
The local provider is a different choice: its execution happens in the application’s environment rather than creating an additional task container. Where host-shell execution is disabled, do not remove that restriction casually. The OpenSandbox review covers sandbox infrastructure separately; a sandbox provider and a complete agent harness solve different layers.
How DeerFlow memory works, and why it still uses tokens
Separate three concerns: the current conversation, reusable memory across conversations, and runtime state needed to resume work. The shared memory configuration distinguishes enabling memory, injecting it into prompts and choosing a backend. “Memory is enabled” therefore does not prove that new facts are being extracted or that every unfinished run survives a restart.
In the reviewed configuration example, DeerMem’s max_injection_tokens belongs under memory.backend_config, with an example value of 2000. Its extraction model is separately configured; omitting the model settings leaves automatic extraction unavailable while non-LLM memory operations can still work. Check the configuration actually loaded by your revision.
Conversation compaction is another mechanism. The summarization middleware provides that layer; it is not the same as durable business memory. Summaries can lose details, and material added to a model’s context still consumes capacity. Do not translate “bounded injection” into “unlimited memory without token costs.”
Our suggested memory test has three parts. Store an innocuous preference, start a new conversation under the same identity, and verify retrieval. Then correct the preference and verify that the old value does not keep winning. Finally, repeat under another test identity and verify that the preference is unavailable there. Inspect logs and storage as well as the final answer.
Record the actual prompt and total provider usage. If memory increases cost without improving accepted outcomes, reduce what is injected or narrow the workflow. If the real problem is a memory layer for an existing agent stack, compare that narrower requirement with our OpenViking memory review rather than replacing the whole stack by default.
Custom DeerFlow skills: procedure, not permission
A useful first skill should encode a repeatable output contract, not an invitation to be generally autonomous. Using the file-based skill format described above, the following is an illustrative SKILL.md body for a source-linked vendor brief. Integrate and test it using your checkout’s skill discovery and approval rules; this is not a tested plugin package.
---
name: vendor-evidence-brief
description: Prepare a source-linked vendor brief for human review.
---
Use only the approved input documents and permitted sources.
Record a source URL or input filename for each factual claim.
Mark missing evidence as unknown; never invent a reference.
Write research branches to separate output files.
Produce a draft, not an approval or a published report.
Do not send messages, change accounts, or execute payments.The distinction matters: “do not send messages” is an instruction, not a technical denial of messaging access. Enforce restricted actions through the tools and credentials available to the run. Treat third-party skill files, bundled scripts and dependencies as code or instructions requiring review, not as automatically trusted additions.
Keep reusable procedure in the skill and task-specific facts in the task inputs. That makes the skill easier to review, compare between revisions and test against both successful and malicious examples. Do not embed live customer secrets in reusable instructions.
A bounded workflow: vendor research to an evidence-backed draft
This is a proposed pilot design, not a reported DeerFlow customer deployment. Start with three approved vendor documents and a fixed question: which options meet a defined technical requirement? Restrict the deliverable to a draft comparison and an evidence file. Do not connect procurement, payment or outbound email actions.
| Stage | Expected artifact | Acceptance check |
|---|---|---|
| Input preparation | Approved files and explicit comparison criteria. | Inputs contain no unnecessary secrets; scope and allowed sources are recorded. |
| Parallel research | A separate evidence file per vendor. | Every factual row has an input reference or permitted source URL; unknowns stay unknown. |
| Synthesis | One draft comparison with limitations. | Claims can be traced to the evidence files; missing evidence is not presented as failure or success. |
| Independent validation | A structured validation result. | Required files exist, machine-readable output parses, and sampled claims match their sources. |
| Human release | An approved or rejected draft. | A named reviewer decides whether anything may be shared or acted upon. |
Include a deliberately conflicting source, a missing document and an interrupted run in the acceptance set. Require the system to surface uncertainty rather than fabricate a clean answer. If a simple single-agent baseline performs better, keep it. Extra subagents are an implementation option, not an outcome metric.
DeerFlow production checklist: what must be demonstrated?
Older “no authentication” descriptions should not be treated as current. The reviewed authentication design includes first-run setup and authenticated user handling. Complete the initial administrator setup through the local /setup flow before making the service reachable by others, and test authorization with separate accounts.
The reviewed production Compose file defaults the published entry point to 127.0.0.1 and requires BETTER_AUTH_SECRET. Preserve loopback exposure during evaluation. Public access should follow an explicit deployment review, not a casual change to BIND_HOST.
Use the following as proposed release gates. They are not a statement that every control is supplied or enabled by DeerFlow.
| Control | Proof to collect |
|---|---|
| Identity and permissions | One user cannot fetch another user’s threads, files or memory; privileged actions require an explicitly authorized identity. |
| Runtime containment | Review mounts, outbound network access, secrets and resource limits. Where AIO uses the host Docker socket, assess that privileged control path separately. |
| Persistence and recovery | Restart during a run, reconnect and inspect state. Verify backup restoration and that retries do not duplicate external actions. |
| Cost and cancellation | Prove spending limits, task deadlines, cancellation behavior and limits on parallel work against representative failure cases. |
| Skill and tool trust | Review executable extensions, MCP launchers and skill changes. Untrusted content must not acquire administrative tool authority. |
| Upgrade ownership | Pin application and image revisions, keep redacted traces, re-run acceptance tests and demonstrate rollback with a named operator. |
For the final release review, use our software QA checklist before launch as a broader delivery checklist, then add the agent-specific cases above. Passing a happy-path demonstration is not enough evidence for account changes, payments or other irreversible actions.
Is DeerFlow worth adopting?
Measure cost per accepted task = (total model costs + total tool costs + total runtime costs + human review costs) / accepted tasks. Use consistent accounting units, count failed attempts within those totals, include memory extraction and delegated calls, and compare the same tasks against your existing baseline. Also record completion time, reviewer corrections, recovery rate and unsafe actions blocked. If no tasks are accepted, report that failure rather than a misleading average.
Choose a DeerFlow pilot when an integrated execution environment might remove real orchestration work. Keep a smaller framework or deterministic workflow when that meets the acceptance contract with less operational complexity. Our LangChain Deep Agents review helps frame the narrower framework option.
Wavect’s AI development service can help scope a bounded workflow, its permissions and acceptance tests. The TwinSoft AI case study is adjacent delivery context, not evidence of a DeerFlow deployment. To assess the fit for your system, discuss the workflow and its failure modes with Wavect.
