AI FLEET
Generate images and voice on a private GPU fleet. Sign in with your email — no password, we'll send you a link.
The link works once and expires in 15 minutes.
AI FLEET · STUDIO
subject + details → style → composition → lighting → quality. Or just write your idea and hit Optimize — it builds this for you.
| Control | Options | Use for |
|---|---|---|
| Model | schnell · dev · sdxl | schnell = fast drafts · dev = best quality · sdxl = stylized |
| Shot | closeup … full body … wide | Framing — the fix for "always a selfie". |
| Aspect | square · portrait · landscape | Shape of the image. |
| Quality | draft · standard · high | Draft to explore, High for the final. |
| Character | you / none | Put your locked look in the shot. |
| Goal | Model · shape | Cues |
|---|---|---|
| Portrait / headshot | dev · portrait | studio lighting, shallow depth of field |
| Product shot | dev · square | seamless white background, soft studio light, sharp |
| Scene / landscape | schnell · landscape | cinematic, golden hour, wide angle, atmospheric |
| Thumbnail / hero | schnell · landscape | bold single subject, high contrast, punchy colors |
Describe the scene, not yourself. Your face, hair, and hat are locked — leave them out and just describe the setting, outfit, action, and lighting. And never ask for text or logos; the model garbles lettering.
Grok (xAI) is fast (~30s); MiniMax Hailuo is higher-detail (~1–3 min, capped ~3/day). It lands in your Library when ready.
| Control | Range | What it does |
|---|---|---|
| Voice | any saved / default | Which cloned voice speaks. "default" = base voice. |
| Speed | 0.5–2.0× | Playback tempo. 0.9–1.1× is most natural. |
| Expressiveness | calm · natural · expressive | Energy & variation in delivery. |
Drop these inside your text where you want the effect. Keep it to 1–3 per sentence.
Free-form works too — try [whisper in a small voice] or [pitch up].
| Vibe | Speed | Expressiveness | Good tags |
|---|---|---|---|
| Hype / announcement | 1.05–1.10 | expressive | [excited] [emphasis] |
| Calm narration | 0.95–1.0 | calm | [low voice] [pause] |
| Dramatic / trailer | 0.90 | natural | [professional broadcast tone] [pause] |
| Funny / casual | 1.05 | expressive | [chuckle] [laughing] |
| Warm / reassuring | 0.95 | calm | [low voice] [short pause] |
| Urgent / alert | 1.15 | natural | [emphasis] [loud] |
| Whispered | 0.95 | natural | [whisper] [low voice] |
MiniMax music-3.0 on the fleet — a full track in ~20–40s. It lands in your Library.
Reference photos keep a specific person, product, or subject consistent across images. Upload one, give it a name, then on the Images tab (or when you ask an agent) say use the reference "<name>" — it works with every image engine.
Private to your account — only you and your own agents can see or use these. (Different from a trained Character — that's a fleet likeness set up separately.)
Everything you and your agents create lands here. Kept 30 days, then deleted.
Everything your agents (via API key / MCP) and you have created on the fleet — newest first.
AI FLEET · CONNECT AN AGENT
Give your AI agent a key and it can make images & voice on your behalf — everything it creates lands right here in your library. Two steps.
On the machine your agent runs on, paste this (it installs the skills and your key — nothing else to set up):
Then refresh your agent. That's it — ask it to "make an image of a sunset" and it shows up in your Library.
Gives your agent (Claude, Cursor, OpenClaw…) tools for image, video, voice, music, text & transcribe. Two image tools: fleet_image (MiniMax — fast, unlimited, subject reference) and fleet_image_openai (OpenAI/ChatGPT — higher detail, transparent background, image-to-image). Just add the block to your MCP config — it downloads the server itself on first launch, and your key is already filled in:
Send the key as a Bearer token to the portal:
Voice is /api/voice/generate. The response returns an image_id/clip_id and a url — GET it with the same Authorization header. List options: /api/images/models, /api/images/characters, /api/voice/voices.
AI FLEET · ADMIN
Who can sign in, and what they can do. Admin = full access + this console · Member = team · Client = studio + their own keys. Roles here override the config list.
Consistent likenesses (LoRAs) placed into any image. Training costs money (fal.ai), so only admins train — then assign a finished character to a client so they can render it (they can't train). A client only ever sees the characters assigned to them.
Everyone who's used the studio, and what they've made. Click a client to view their files.
Every client the router has served, by hostname/IP — whether the agent used the MCP or hit the fleet-router API directly. Portal/MCP traffic auto-attributes by account email; map the rest to a person here.
AI FLEET · GUIDE
Welcome. This studio lets you create images and voice audio using our own AI, running on private hardware. No software to install — it all works right here in your browser.
There's no password. Enter your email and we send you a link — click it and you're in. The link works once and lasts 15 minutes. You stay signed in on this device afterwards.
You describe what you want, choose a few options, and press Generate. It takes a few seconds to about a minute. When it's done it appears below the buttons, and it's saved to your Library automatically.
Everything you create is private to you and kept for 30 days, then removed automatically. Download anything you want to keep.
There are two ways to make an image. Use whichever feels easier.
me hiking a mountain trail at sunrise.Type your description in the Prompt box. Describe one clear subject and a setting: a red fox in fresh snow, golden evening light works far better than just fox. Open Image reference at the bottom for clickable words you can add (styles, lighting, framing).
AI assist: a cozy coffee shop on a rainy evening → Optimize → Generate. Then click the image to expand it and see the full details.
Type what you want said, pick a voice, and press Generate speech. You'll get an audio clip you can play and download.
You can drop little tags right into your text to shape the delivery — things like [chuckle], [whisper], [excited], [pause]. Open Voice reference at the bottom and click any tag to insert it. Keep it to one to three per sentence.
[professional broadcast tone] Big news. [pause] [excited] We just launched! [emphasis] Try it now.
Not sure how to phrase it? In the AI assist box, describe what you want — a calm, slow welcome message for a spa — and press ✦ Craft the script. It writes the words and sets the speed and expressiveness for you.
A character is a saved likeness — a person the AI can place into any image so they look consistent every time. If you have one set up, pick it from the Character dropdown on the Images tab.
Describe the scene, not the person. Your face, hair, and hat are already locked in — so you only need to describe where you are and what you're doing. Say on stage at a conference, dramatic lighting, not bald man with a beard on stage. Describing your own features actually makes the result worse.
The AI assist handles this automatically — when a character is selected it writes a scene-only prompt for you.
The Library tab holds everything you've made — only ever your own work.
You can give an AI agent — Claude, Cursor, OpenClaw, anything that speaks MCP — direct tools to make images, video, voice, music, chat and transcribe across the whole fleet. Everything it creates lands in your library, exactly like the studio.
Open 🔑 API keys at the top, create a key, and copy it. Your ready-to-paste command and MCP config appear there with the key already filled in — grab it from there and drop it into Step 2.
⚠ Treat your key like a password. Never paste it in a shared or public channel (Discord, Slack, a group chat). If a key ever gets exposed, revoke it on the API keys page and create a new one — it takes ten seconds.
sh isn't available)? Download it yourself:
"command": "python3", "args": ["/home/YOU/fleet_mcp.py"] with the full absolute path to the downloaded file (never a placeholder like /path/to/fleet_mcp.py).Restart your agent and it now has fleet tools. Ask it to "make an image of a sunset" or "read this out in a calm voice" and the result appears in your Library.
fleet_image (MiniMax, fast & unlimited, subject-reference) and fleet_image_openai (OpenAI/ChatGPT subscription — higher detail, transparent background, image-to-image).fleet_usage shows what's left on every subscription (and when each resets), so your agent won't blow a cap.The server ships with usage guidance (which tool to pick, which models are free vs. frontier, and to check caps before capped work) and multi-step recipes — ready-made playbooks like image → video, narrated clip, consistent character, and transcribe → summarize. Compatible agents load these automatically, so you can just ask for the outcome ("turn this photo into a short video") and it chains the tools for you.
minimax-m3 is a chat model — you talk to it for text, and it runs on the fleet's language router (your own GPUs + Ollama Cloud passthrough). It counts as Ollama activity and isn't tied to any media cap.
Images / video / voice / music go directly to the MiniMax subscription API — those are the image-01, MiniMax-Hailuo-02, speech-02-hd and music-1.5 models. They count against your MiniMax plan caps (e.g. video is limited per day). Same brand, different service, different meter.
[pause] and normal punctuation for natural pacing.