Planner + ReAct multi-agent architecture ยท A2A & MCP native ยท Sandboxed execution ยท One-command deploy
English ยท ็ฎไฝไธญๆ
MultiGen is an open-source, general-purpose AI Agent platform designed for fully private, on-premise deployment. It pairs a Planner agent (decomposes user goals into steps) with a ReAct agent (executes each step using tools), and runs every action inside an isolated Docker sandbox โ so your data never leaves your infrastructure.
Out of the box, MultiGen can browse the web, run shell commands, generate images / videos / 3D models / TTS audio, build slide decks and reports, and orchestrate other agents via A2A and external tools via MCP.
๐ก Think of it as your private, self-hosted alternative to Manus / Claude Agent / GPT Agent โ but you own the data, the model, and the stack.
MultiGen ships two long-lived branches โ pick the one that matches your scenario:
Scenario Branch Use it for ๐ฅ๏ธ Local Docker deployment masterLocal one-command Docker stack, evaluation, development, contributing ๐ Online / production deployment onlinePublic / production environments โ battle-tested, with hotfixes & deployment configs verified online Local Docker (this branch โ
master):# ๐ฅ๏ธ Local Docker deployment โ use master git clone https://github.com/LiXiaoYaoCareFree/MultiGen.git cd MultiGen docker compose up -d --buildOnline / production:
# ๐ Online / production deployment โ use online git clone -b online https://github.com/LiXiaoYaoCareFree/MultiGen.git cd MultiGen docker compose up -d --build
โ ๏ธ Never deploymasterto a public / production environment โ onlyonlineis verified for that. Keep production in sync by pulling fromonlineonly.
| ๐ง Planner + ReAct architecture | A two-stage agent: the Planner breaks down the goal into JSON sub-steps, the ReAct agent iteratively reasons & acts on each step. |
| ๐ MCP & A2A native | Plug in any MCP server (search, maps, code, custom tools) and delegate sub-tasks to peer agents via Agent-to-Agent protocol. |
| ๐ก๏ธ Sandboxed execution | Every shell / browser / file action runs inside an isolated Ubuntu + Chrome + VNC container. The model can't touch your host. |
| ๐จ Multimodal generation | Built-in tools for image (Volcengine / SD), video, 3D models, TTS (Qwen / podcasts), virtual anchors, audio mixing, slide decks. |
| ๐ Any OpenAI-compatible LLM | Works with DeepSeek, Volcengine, SiliconFlow, Qwen, OpenAI, vLLM, Ollama, etc. โ just edit config.yaml. |
| ๐ข One-command deploy | docker compose up -d --build brings up the full stack: UI, API, sandbox, Postgres, Redis, Nginx. |
| ๐ก Real-time streaming UI | SSE-driven Next.js frontend renders plans, tool calls, intermediate results, and final answers live. |
| ๐ Replayable sessions | Full session state in PostgreSQL; generated files mirrored locally and to Tencent COS for replay & sharing. |
One screen, three layers of MultiGen at work: persistent session history, a live Planner+ReAct execution stream, and the agent's sandbox computer rendering the paper in real time.
The screenshot above captures MultiGen tackling a real task โ "Analyze the AI-Researcher: Autonomous Scientific Innovation paper at alphaxiv.org/abs/2505.18705" โ and showcases three of the platform's most distinctive capabilities in a single view:
Every conversation is a fully replayable session, stored in PostgreSQL and synced to Tencent COS. The sidebar in the screenshot shows the breadth of tasks MultiGen handles out of the box:
| Visible session | Tools exercised |
|---|---|
| ๐ Baidu tech-ops weekly charts | browser ยท file ยท shell |
| ๐ป GitHub Java project discovery | search ยท browser |
| ๐งฎ SQLite + FAISS data vectorization | shell ยท file |
| ๐ PDF batch download & merge from GitHub | browser ยท file ยท shell |
| ๐ฏ Late-autumn Hangzhou Faming Temple image search | search ยท image_generation |
| ๐ AI-Researcher paper reading (active) | browser ยท file ยท mcp |
| ๐๏ธ Article voice-over + song audio mixing | qwen_tts ยท audio_mixing |
| ๐ง 3D pet model retrieval & rendering | model_3d ยท browser |
๐งช autoresearcher / AI-Scientist / sibyl-research-system paper deep-dives |
browser ยท file ยท a2a |
| ๐ฌ Automated cute-video generation pipeline | volcano_image ยท volcano_video ยท video_concatenation ยท virtual_anchor |
Sessions persist across restarts and can be reopened, branched, or replayed step-by-step โ powered by SQLAlchemy async + Alembic migrations.
The center column streams the agent's reasoning in real time over SSE. For this task you can see the two-stage architecture cleanly:
- PlannerAgent parses the user goal and emits a JSON plan โ fetch URL โ browse page โ download PDF โ extract content โ summarize.
- ReActAgent picks up each step and iteratively reasons โ calls a tool โ observes the result โ continues:
- โ
่ฎฟ้ฎ่ฎบๆ้พๆฅโbrowser.goto(https://www.alphaxiv.org/abs/2505.18705) - โ
ๆญฃๅจๆๅผ็ฝ้กตโbrowser.snapshot()returning the page DOM - โ
ๆญฃๅจๆต่ง็ฝ้กตโ extracting title, abstract, sections - โ
ๆญฃๅจไธ่ฝฝๆไปถโbrowser.download()of the PDF - โ
ๆญฃๅจๆๅผๆไปถโfile.read(.../2505.18705.pdf)to ingest content - ๐ ...continues until the ReAct loop summarizes the paper
- โ
Every green check is a discriminated event (plan ยท step ยท tool ยท message ยท done) flowing through /api/sessions/{id}/chat โ defined in api/app/domain/models/event.py and produced by PlannerReActFlow in api/app/domain/services/flows/planner_react.py.
The right pane is not a static screenshot โ it's a live window into the agent's isolated Docker sandbox. As the ReAct loop drives the headless Chrome inside the sandbox (Ubuntu + Chrome + VNC, port 8080), you see exactly what the agent sees:
- ๐ The alphaxiv.org paper rendered inside the sandbox browser
- ๐ The PDF preview with "Highlight of Key Insights" section in view
- ๐ Scroll / click / extract events mirrored frame-by-frame
This is full "computer use" transparency โ your model can browse, click, type, and download, but it's all firewalled inside a disposable container. Your host machine is never touched, and every action is observable and auditable.
๐ก๏ธ Why this matters for private deployment: the model never gets a shell on your infrastructure. Every
shell,browser, andfiletool call is proxied to the sandbox container, which can be destroyed and rebuilt at will.
Agent plans the search, calls the right tool, and synthesizes a sourced answer.
Image search and ranking, with live previews streamed back to the UI.
LLM provider โ connect any OpenAI-compatible endpoint |
Agent behavior โ iterations, retries, search depth |
MCP servers โ plug in external tools live |
A2A agents โ federate with peer agents |
End-to-end creative workflow โ from prompt, to plan, to rendered assets.
Generated portrait #1 |
Generated portrait #2 |
Generated portrait #3 |
Generate full multi-speaker podcasts with Qwen-TTS, automatically mixed with background music.
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Next.js UI (3000) โ
โ Plans ยท Steps ยท Tool calls ยท SSE stream โ
โโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโ
โ /api (SSE)
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ FastAPI (8000) โ
โ โโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโโ โ
โ โ AgentService โ โ โ AgentTaskRunner โ โ
โ โโโโโโโโโโโโโโโโ โโโโโโโโโโฌโโโโโโโโโโโ โ
โ โผ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ PlannerReAct Flow โ โ
โ โ Planner โโบ ReAct (loop) โ โ
โ โโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ tools โ
โ โโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโโ โ
โ โ file ยท shell ยท browser ยท search ยท MCP โ โ
โ โ image ยท video ยท 3D ยท TTS ยท A2A ยท ... โ โ
โ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โ
โโโโโโโฌโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโฌโโโโโโ
โผ โผ โผ
โโโโโโโโโโโโ โโโโโโโโโโโโ โโโโโโโโโโโโโโโโโโโโ
โPostgreSQLโ โ Redis โ โ Docker Sandbox โ
โ sessions โ โ streams โ โ Ubuntu + Chrome โ
โโโโโโโโโโโโ โโโโโโโโโโโโ โ + VNC (8080) โ
โโโโโโโโโโโโโโโโโโโโ
Agent execution flow:
AgentServicereceives a chat message โ dispatches it to anAgentTaskRunnervia Redis Streams.AgentTaskRunnerrunsPlannerReActFlow:- PlannerAgent โ decomposes the request into a JSON plan of sub-steps.
- ReActAgent โ for each step, iteratively reasons โ calls a tool โ observes โ continues, then summarizes.
- Events stream back via SSE (
planยทtitleยทstepยทmessageยทtoolยทwaitยทerrorยทdone).
- ๐ณ Docker
>= 20.10 - ๐ Docker Compose
>= 2.0 - ๐ An API key for any OpenAI-compatible LLM (DeepSeek / Volcengine / OpenAI / vLLM / Ollamaโฆ)
๐ก Pick the right branch for your deployment scenario:
- ๐ฅ๏ธ Local Docker deployment โ use
master(this branch)- ๐ Online / production deployment โ use
online
# ๐ฅ๏ธ Local Docker deployment โ use master (default branch)
git clone https://github.com/LiXiaoYaoCareFree/MultiGen.git
cd MultiGen
# ๐ Online / production deployment โ use online instead
# git clone -b online https://github.com/LiXiaoYaoCareFree/MultiGen.git
# cd MultiGenCreate a .env file in the project root:
# โโ Required โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
COS_SECRET_ID=your_cos_secret_id_here # Tencent COS SecretId
COS_SECRET_KEY=your_cos_secret_key_here # Tencent COS SecretKey
COS_BUCKET=your_cos_bucket_here # COS bucket name
OPENAI_API_KEY=your_llm_api_key_here # LLM API key
# โโ Optional โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
NGINX_PORT=8088 # public port
ADMIN_API_KEY=your_admin_api_key_here # admin auth key
LLM_PROVIDER=volcano # deepseek / openai / volcano
TENCENT_AI3D_API_KEY=... # for 3D model generation
DASHSCOPE_API_KEY=... # for Qwen-TTSEdit api/config.yaml:
llm_config:
base_url: https://api.deepseek.com/
api_key: YOUR_DEEPSEEK_API_KEY
model_name: deepseek-reasoner
temperature: 0.7
max_tokens: 8192
agent_config:
max_iterations: 100
max_retries: 3
max_search_results: 10
mcp_config:
mcpServers:
amap-maps-streamableHTTP:
transport: streamable_http
enabled: true
url: https://mcp.amap.com/mcp?key=YOUR_AMAP_API_KEY
jina-mcp-server:
transport: streamable_http
enabled: true
url: https://mcp.jina.ai/v1
headers:
Authorization: Bearer YOUR_JINA_API_KEYdocker compose up -d --buildVisit http://localhost:8088 (or whichever NGINX_PORT you set). The API health probe lives at /api/status.
| Tool | Purpose |
|---|---|
file |
Read / write / patch files inside the sandbox |
shell |
Run shell commands in the sandbox |
browser |
Headless Chrome โ navigate, click, extract, screenshot |
search |
Web search (Bing / Google / Jina) |
message |
Ask the user a clarifying question mid-task |
image_generation ยท volcano_image |
Text-to-image generation |
volcano_video ยท video_concatenation |
Text-to-video & post-processing |
model_3d |
Text/image-to-3D via Tencent AI3D |
virtual_anchor |
Avatar / digital-human video |
qwen_tts ยท audio_mixing |
TTS + multi-track audio mixing |
mcp |
Call any registered MCP server |
a2a |
Delegate a sub-task to a peer agent |
๐ To add your own tool, see CLAUDE.md โ Adding a New Tool.
MultiGen/
โโโ api/ # Backend API service (FastAPI)
โ โโโ app/ # Domain / application / infrastructure layers
โ โโโ tests/ # Pytest suite
โ โโโ config.yaml # Runtime LLM / MCP / A2A config
โโโ ui/ # Frontend (Next.js 14, App Router)
โโโ sandbox/ # Sandbox runtime (Ubuntu + Chrome + VNC)
โโโ nginx/ # Reverse-proxy gateway
โ โโโ nginx.conf
โ โโโ conf.d/default.conf
โโโ assets/ # Screenshots used in this README
โโโ docker-compose.yml
โโโ .env # Environment variables (create your own)
โโโ README.md
| Container | Service | Description |
|---|---|---|
manus-nginx |
Nginx | Reverse-proxy gateway, the only exposed entrypoint |
manus-ui |
Next.js | Frontend UI |
manus-api |
FastAPI | Backend API |
manus-postgres |
PostgreSQL | Session & message store |
manus-redis |
Redis | Task streams & cache |
manus-sandbox |
Sandbox | Ubuntu + Chrome + VNC isolated runtime |
# Start everything (detached) + rebuild images
docker compose up -d --build
# Check service status
docker compose ps
# Follow logs
docker compose logs -f
docker compose logs -f manus-api
docker compose logs -f manus-ui
# Restart a single service
docker compose restart manus-api
# Stop everything
docker compose down
# Stop and wipe data volumes (DANGEROUS โ deletes the database)
docker compose down -v- Place your TLS files in
nginx/ssl/:fullchain.pemprivkey.pem
- In
nginx/conf.d/default.conf, add/enable alisten 443 sslserver block pointing at those files. - In
docker-compose.yml, enable the443:443port mapping (and mountnginx/sslif needed). - Apply changes:
docker compose restart manus-nginx
Each sub-project has its own dev guide:
- ๐ง API service โ FastAPI, SQLAlchemy async, Alembic, Pytest
- ๐จ UI service โ Next.js 14, App Router, SSE streaming
- ๐ฆ Sandbox service โ Ubuntu + Chrome + VNC runtime
Quickstart for the API:
cd api
python -m venv .venv && source .venv/bin/activate
pip install uv && uv pip install -r requirements.txt
playwright install
uvicorn app.main:app --host 0.0.0.0 --port 8000 --reload- Planner + ReAct dual-agent flow
- MCP & A2A integrations
- Multimodal tools (image / video / 3D / TTS)
- DeepSeek reasoning-model (v4) compatibility
- Long-term memory / RAG plugin
- Multi-user workspace permissions
- Plugin marketplace for tools & MCP servers
- Mobile-friendly UI
Contributions are warmly welcomed โ issues, PRs, tool plugins, and translations alike.
- Fork the repository
- Create your feature branch (
git checkout -b feat/amazing-thing) - Commit your changes (
git commit -m 'feat: add amazing thing') - Push to the branch (
git push origin feat/amazing-thing) - Open a Pull Request
Please read CLAUDE.md first โ it documents the architecture, the agent contracts, and how to add new tools / LLM providers safely.
MultiGen stands on the shoulders of these excellent projects:
- FastAPI ยท Next.js ยท SQLAlchemy
- Model Context Protocol ยท A2A
- Playwright ยท Docker
- DeepSeek ยท Volcengine ยท SiliconFlow ยท Qwen โ for outstanding open-source LLM endpoints
Released under the MIT License.
If MultiGen is useful to you, please consider giving it a โญ โ it really helps!
Made with โค๏ธ for builders of private AI agents.







