# What's New ?

We've rebuilt Qolaba from the ground up, a cleaner home, a smarter sidebar, refreshed chat and unified studios for image, video and audio. Every interaction now feels faster, simpler and more intuitive, from your first prompt to your final output.

### 1. Reimagined Home

A cleaner dashboard that puts recent work and every tool one click away.

<figure><img src="/files/TUV6RkUBx9TfrIyEeF9I" alt=""><figcaption></figcaption></figure>

### 2. A new sidebar

Redesigned navigation to move between Home, Tools, Context and account in one tap.

<figure><img src="/files/FOZqZ9YkV0BraBPEa8tS" alt=""><figcaption></figcaption></figure>

### 3. Refreshed chat

A new composer, model picker and a calmer conversation view.

<figure><img src="/files/lXrGlB4wzRegnRj4bgQh" alt=""><figcaption></figcaption></figure>

### 4. Rebuilt studios

Image, video and audio now share one consistent, focused layout.

<figure><img src="/files/DpeJZDJVJoVfV2mKFWcX" alt=""><figcaption></figcaption></figure>

<figure><img src="/files/x7PkqbOckoKknectv3JE" alt=""><figcaption></figcaption></figure>

### 5. Your Context, All at one Place

Knowledge bases, files, agents, and prompts now have their own dedicated workspace to organize, search, and reuse across your workflows.

<figure><img src="/files/HGx6xUU1ghqC7nnocLuE" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
**Your data is 100% intact.** Nothing was deleted or reset. Every chat conversation, generated image, video, audio file, agent, knowledge base, saved prompt, uploaded file, credit balance, and team workspace has been carried over exactly as it was. API keys and integrations continue to work without any changes. All existing share links still resolve correctly.
{% endhint %}


# Navigation

The most visible change in the new Qolaba is the left navigation rail. In the old design, the sidebar listed eight items — many of which were top-level labels for tools that each had their own sub-pages (like **Create Images** expanding to Text To Image, Image To Image, and Image Editing). This made the nav feel long and required multiple clicks to reach a tool.

The new rail has just six items: **Home, Chat, Image, Video, Audio,** and **Context**. Each one takes you directly to a unified workspace — no sub-pages to expand. Tools that previously lived under *Audio & Video* are now mode cards inside a single Audio Studio or Video Generator page.

The old Dashboard (with its welcome message and tool tiles) has become the **Home** page. The tile grid is still there — scroll down past the prompt hero and you'll find the **Apps** section with quick links to every tool.

### Before

![Old Qolaba dashboard showing long left nav with Dashboard, History, Write & Chat, Create Images with sub-items, Audio & Video with sub-items, Discover, Credits & Billing, Refer & Earn. Main area shows welcome message, Generate Image/Video/Speech tiles, Qommunity banner.](/files/dd617728c4a4253301e53c7d17da62b174940a03)

The old Dashboard had 8 nav items, several with expandable sub-menus. The main area was a landing page with shortcut tiles and a community banner.

### After

![New Qolaba Home showing compact left rail with Home, Chat, Image, Video, Audio, Context. Workspace block top-left shows org name and credits. Main area shows personalized greeting and Continue where you left off section.](/files/f3c9b81c36a39b7990d89c8c521e1636e2100b18)

The new Home is a *starting point*, not just a dashboard. Type anything in the prompt box and Qolaba routes you to the right tool automatically. Your recent work appears below so you can pick up exactly where you left off.

![Home page scrolled down showing APPS section with Chat, Image, Video, Audio, and History tiles under the heading 'Everything in one place'](/files/d711bce849a3c9ebc8540341d473fdddc3cb6902)

Scroll down on Home to reach the **Apps** grid. This is where the old Dashboard's shortcut tiles now live — a direct link to every tool from one place.

| Was                                        | Now — and why it changed                                                                                             |
| ------------------------------------------ | -------------------------------------------------------------------------------------------------------------------- |
| 8-item left nav with collapsible sub-menus | **6-item rail, no sub-menus.** Each item is a direct destination — fewer clicks to reach any tool.                   |
| Dashboard with welcome message + tile grid | **Home with smart prompt + Apps grid below fold.** You can now start any task from Home without navigating first.    |
| History as a sidebar nav item              | **Recent list in the sidebar** + See all link. Your last few items are always one glance away.                       |
| Credits & Billing buried in the nav        | **Credits shown in the workspace block** at the top of the sidebar — visible at all times without navigating away.   |
| Org switcher in the top-right profile menu | **Org + workspace selector at the top of the sidebar** — easier to see which workspace you're in and switch quickly. |


# History

Previously, History lived as its own page in the sidebar. You'd click it, land on a tabbed page, and filter by All / Images / Audio / Video / Favourites. It was separate from your current context.

Now your recent work surfaces *in-context*. The sidebar always shows your last several items with type icons (chat bubble, image, video, audio waveform), so you can spot and resume a recent session instantly. When you need the full list, hit **See all** for the complete history view with all the same filters you had before.

### Before

![Old History page showing tabbed interface with All, Images, Audio, Video, Favourites tabs and a masonry grid of generated items with star and menu icons](/files/552792f1b699f6660d3b362005d76b378bae61ba)

History was a standalone page — you had to navigate away from your current work to look something up.

### After

![Home page showing sidebar Recent section with pinned items and a History tile in the Apps grid](/files/d711bce849a3c9ebc8540341d473fdddc3cb6902)

Recent items live in the sidebar, always visible. The History tile in the Apps grid opens the full paginated view when you need to search or bulk-manage older items.

| Was                                            | Now — and why it changed                                                                      |
| ---------------------------------------------- | --------------------------------------------------------------------------------------------- |
| Sidebar nav → History → tabbed full-page view  | **Sidebar Recent list** (always visible, no navigation needed) + **See all** for full history |
| All / Images / Audio / Video / Favourites tabs | Same filters available in the full history view                                               |
| Star to favourite, 3-dot menu for actions      | Same — pinned items shown with a 📌 badge at the top of the Recent list                       |


# Chatbot

The Chat experience itself hasn't changed in any fundamental way — you can still pick any model, use agents, run the toolkit, branch conversations, and generate images inline. What changed is *how you start* a chat and *where context lives*.

In the old UI, everything was inside the chat screen itself: a New Chat button at the top-left, and sidebar tabs (History, Knowledge, Files, Prompts) that you'd open to load context into your conversation. This made the chat UI feel crowded, especially on smaller screens.

In the new design, starting a chat is even simpler — just type in the Home prompt box and you're in. Context (knowledge bases, files, saved prompts) has moved to its own dedicated section called **Context** in the left rail, giving it more space and better organisation than sidebar tabs ever could.

### Before

![Old Write & Chat page showing left sidebar with New Chat button, History/Knowledge/Files/Prompts tabs, and main area with How can I help you today prompt and model/agent/toolkit row](/files/f15911071e0487da1c90f14fd9a89f984ec199b1)

Everything was in one place: nav, context tabs, and the chat itself. Useful, but the sidebar tabs were small and the chat area felt constrained.

### After

![New Chat empty state leading with agent cards: General Assistant, Copywriter, Code Helper, Create your own, and Browse all agents link below](/files/571ceb5a3c31fc62daf6bb6b757107a640630179)

The new chat empty-state surfaces agents upfront — the most common starting point for structured tasks. The prompt box is at the bottom; context is managed separately in the Context section.

<table><thead><tr><th width="403.63671875">Was</th><th>Now — and why it changed</th></tr></thead><tbody><tr><td><strong>New Chat</strong> button in sidebar</td><td>Type on Home, click <strong>＋</strong> in the rail, or <code>⌘⇧O</code> — start from anywhere</td></tr><tr><td>Agents hidden in a small prompt-bar dropdown</td><td>Agents lead the empty state; same dropdown available mid-conversation</td></tr><tr><td>Toolkit as a single prompt-bar toggle</td><td>Toolkit side panel — all 12 tools visible and individually toggleable</td></tr><tr><td>Sidebar tabs: History, Knowledge, Files, Prompts</td><td>Moved to <strong>Context</strong> in the left rail (see section 8)</td></tr><tr><td>Model selector, Chat Branching, Thinking Depth, Temperature</td><td>All unchanged — same location, same controls</td></tr></tbody></table>


# Starting a New Chat

The **New Chat** button is gone. Instead there are three equivalent ways to start fresh, depending on where you are:

* **From Home** — just type in the prompt box. Qolaba opens a new chat automatically.
* **From anywhere** — click the **＋** icon next to Chat in the left rail.
* **Keyboard shortcut** — `⌘⇧O` on Mac, `Ctrl⇧O` on Windows/Linux (unchanged).

#### Before

![Orange New Chat button in the top left of the old chat sidebar](/files/f15911071e0487da1c90f14fd9a89f984ec199b1)

Old: a dedicated button — easy to find, but it meant navigating to the Chat page first.

#### After

![New chat — no standalone New Chat button, but the Home prompt box and the + icon in the sidebar both start new chats](/files/571ceb5a3c31fc62daf6bb6b757107a640630179)

New: start a chat from Home (most common) or tap + in the rail. You never need to be "inside" Chat to begin a conversation.


# Agents

Agents are custom AI personas or pre-built assistants that specialise in a task — writing, coding, research, customer support, and so on. In the old design, you selected an agent from a small dropdown inside the prompt bar, which made them easy to miss.

In the new design, **agents lead the empty state**. When you open Chat, the first thing you see is a set of featured agents. This makes it obvious that agents are the right starting point for structured or repeatable work. You can also browse the full catalogue and create custom agents — all from the same screen.

#### Before

![Old chat prompt bar showing Personal Assist... agent label inline next to model selector — easy to overlook](/files/f15911071e0487da1c90f14fd9a89f984ec199b1)

Old: agent selection was a small inline dropdown in the prompt bar. Most users never noticed it was there.

#### After

![New chat agent selector dropdown open mid-conversation showing General Assistant, Copywriter, Code Helper and custom agent options](/files/0903084256dda4b2065c515ac38ea59c40e6cbce)

New: agents are prominently featured on the empty state. Mid-conversation, the same dropdown is still available in the prompt bar — pick a different agent at any point to shift focus.

![Browse agents screen showing a grid of pre-built agent templates with names and description previews](/files/592da048941e412a57579c01c00e950e56a8ce3b)

**Browse all agents** opens a full catalogue of pre-built templates across categories like Writing, Coding, Research, and Business. Click any to start a conversation immediately, or duplicate one as the basis for a custom agent.


# Toolkit

The Toolkit is a set of 12 optional capabilities — web search, image generation, file reading, code execution, text-to-speech, and more — that you can switch on or off per conversation. In the old UI it was a small toggle in the prompt bar row, which gave no indication of what was available.

It now opens as a dedicated **side panel** with labelled toggles for every tool. The Auto/Manual mode setting (Auto = Qolaba decides when to use tools; Manual = you control) is in the same place.

#### Before

![Old chat prompt bar row showing Tool Kit label as a small icon-label toggle at the right of the bar](/files/f15911071e0487da1c90f14fd9a89f984ec199b1)

Old: a compact toggle with no visibility into what tools were available or active.

#### After

![New Toolkit side panel showing 12 individually labelled toggles including Web Search, Image, Video, URL, File, Run Code, Spreadsheet, STT, TTS, Music, PII Protection in Auto mode](/files/aeaac484f9784ba2a4cc4c286b6fbf17620a89b3)

New: every tool is visible and individually toggleable. You can see at a glance which capabilities are active, and switch Auto/Manual without hunting for a setting.


# Chat Branching

Branching lets you explore multiple directions from any point in a conversation, like forking a decision tree. Ask a question, branch at the answer, try different follow-ups, compare results side by side. This is one of Qolaba's most powerful features for creative and analytical work.

![Chat Branching view showing a conversation tree with the original thread on the left and multiple branch paths stemming from a decision point](/files/8f809a13433b427d81f4c9a451042c9d8b2d8d06)

**Chat Branching is unchanged.** Open it with `⌘⇧B` (Mac) or `Ctrl⇧B` (Win/Linux). The tree view, branch creation, and merge controls work exactly as before.


# Context

In the old Qolaba, the things you'd bring into a conversation — knowledge bases, uploaded files, saved prompts — lived as tabs inside the Chat sidebar. This was practical for a single session, but it had real limitations: you couldn't search or filter your knowledge bases, you couldn't see which files were part of which KB, and your saved prompts were buried in a tab that most users never opened.

All of this has been promoted to its own section in the left rail: **Context**. Think of it as your personal AI library. It has four tabs — **Agents, Knowledge bases, Files, Prompts** — each with a full-page view, search, sorting, and filtering. Everything you saved in the old system is here, unchanged.

### Knowledge Bases

A knowledge base is a collection of documents, URLs, or files that you want your AI to draw on during a conversation. You attach a KB to any chat, and Qolaba will reference its contents when answering your questions. Previously you could only see KBs as a flat list in a sidebar tab. Now each KB shows its source count, its sync status (Ready / Processing / Error), and when it was last updated.

#### Before

![Old Chat sidebar Knowledge tab showing + Create button, Me filter toggle, German Campaign (0 Files) and Optimization Docs (3 Files) as checkbox items with Delete and Add To Chat buttons at the bottom](/files/5c86d8f9779c3c15f545347e3271fed4cb708333)

KBs as a simple checkbox list inside the chat sidebar. You could attach them to the current chat, but there was no way to manage, search, or see their status without opening each one.

#### After

![New Context → Knowledge bases page showing 3 KB cards: Brand Guidelines (3 sources, Ready), Graphic assets (5 sources, Ready), Optimization docs (3 sources, Ready). Plus New knowledge base card on left. Search and filter controls at top right.](/files/263c6b64fcf64ff02799808c1663a8ffb15418a1)

KBs now have a dedicated page with card layout. Each card shows source count and a colour-coded status badge. You can search, filter by status (Ready / Processing), and create new KBs — all without being inside a chat.

### Files

Files are the raw uploads that power your knowledge bases or that you drop into individual chats for reference. The old chat sidebar just showed a flat scrollable list of filenames with delete buttons. With 197 files, that list was impossible to navigate.

The new Files view is a proper **data table**: filename, which KB it belongs to (if any), file size, and when it was added. You can sort by any column, filter by file type or KB membership, and upload new files from the same page.

#### Before

![Old Chat sidebar Files tab showing a flat scrollable list of filenames (credit optimization.p..., Credit Insights Finalized..., Tool Credits Reference-20... etc.) with trash/delete icons on the right](/files/be2f7efa40105bef3074baf42b3997841259d8ea)

A plain scrollable list. No sorting, no filtering, no way to see file sizes or which KB they belonged to. Finding a specific file meant scrolling.

#### After

![New Context → Files page showing a sortable table with columns: Name (with file type icon), In Knowledge Base, Size, Added. 197 files listed. Sort controls for Newest, Type, Knowledge base at top right. Upload files button at top right.](/files/94a4295f0c135da25b112b7ea967957ecafca452)

A full data table with sortable columns. The *In Knowledge Base* column immediately shows which files are organised into KBs and which are loose uploads — critical for managing a large library.

### Saved Prompts

Saved prompts are reusable templates — crafted instructions that you want to use across multiple conversations without retyping them. Examples: a system prompt that defines a persona, a structured research framework, a creative writing brief. In the old UI they were cards in a chat sidebar tab, with copy/fork/edit/delete actions.

In the new Context, prompts have their own tab with a cleaner card grid. Each card has a **Use in chat** button (opens a new chat with the prompt pre-filled) and **Copy**. The management actions (edit, delete) are behind the `···` menu to keep the cards uncluttered.

#### Before

![Old Chat sidebar Prompts tab showing + New Prompt button, and five saved prompt cards with title, preview text, and icon row showing copy/fork/edit/delete actions](/files/acb3cbec12062a970b5c1e6e53e2362938401b53)

Prompts were accessible but only from inside a chat session. Adding, editing, or organising them was awkward because you had to have a chat open first.

#### After

![New Context → Prompts tab showing 5 prompt cards in a grid layout, each with title (What is paper trading, What is trading psychology, etc.), preview text, Use in chat arrow button, and Copy button](/files/db22d409088f98b8f2187a31ea9202c539784dd2)

Prompts are now managed independently of any chat. The **Use in chat →** button is the main action — tap it and a new conversation opens with the prompt pre-loaded. You can manage your prompt library at any time, even between sessions.

### Agents (new in Context)

Context also introduces a new **Agents** tab — a dedicated space to browse pre-built agents and manage the custom agents you've created. This didn't exist as a standalone page before; agents could only be accessed from inside the Chat interface. Now you can plan, edit, and organise your agents independently.

![New Context → Agents tab showing a grid of agent cards including pre-built agents (General Assistant, Copywriter, Code Helper) and user-created custom agents with edit and share options](/files/c20a1ed0b22b199e308b073b1acf8ad4e0d7f10d)

**Context → Agents** is a new addition. Browse the full catalogue of pre-built agents, see your custom agents, edit their system prompts, adjust their toolkit access, and manage sharing permissions — all from one page.

| Was                                                    | Now — and why it changed                                                                |
| ------------------------------------------------------ | --------------------------------------------------------------------------------------- |
| Chat sidebar tab → Knowledge (flat KB list)            | Left rail → **Context → Knowledge bases** — card view with status, source count, search |
| Chat sidebar tab → Files (scrollable list)             | Left rail → **Context → Files** — sortable table with KB membership, size, date added   |
| Chat sidebar tab → Prompts (cards with copy/fork/edit) | Left rail → **Context → Prompts** — same prompts, with prominent *Use in chat* action   |
| No standalone Agents page                              | Left rail → **Context → Agents** (new) — manage all agents independently of any chat    |
| Attaching a KB required being inside a chat            | KBs can be managed from Context; attach to any chat via **＋** in the prompt box         |


# Image Generation

Image generation had three separate pages in the old UI: **Text To Image** (generate from a prompt), **Image To Image** (use a reference), and **Image Editing** (inpaint, upscale, remove background, create variations). Switching between these modes required clicking different nav items, losing your place in the process.

All three are now unified in a single **AI Image Generator** workspace. You start by picking a model card or a task chip — Qolaba shows you the right controls for that task automatically. Reference images, editing tools, and generation all coexist in the same canvas.

### Before

![Old Create Images page showing Text To Image active, with model dropdown (Nano Banana 2), select preset, private session toggle, prompt box, Magic Prompt, Keywords, Negative Keywords in the left panel. Right panel shows No. of Generations 1-4, Quality 0.5K-4K, and Dimension grid.](/files/7ce9628d10f8e52a1192478139e46b71c53068c1)

Three separate sidebar items for three modes. Switching from text-to-image to image-to-image meant navigating away and starting fresh.

### After

![New AI Image Generator showing model cards (Flux 1.1 Pro, Gemini Nano Banana, GPT Image, Ideogram) and task chips (Photoreal shot, Illustration & art, Logo & text, From a reference image). Right panel shows Number of images 1-4, Aspect ratio grid, Quality, Advanced.](/files/3e33edcf7ac38b5fb7e780efcce7a9f3dfaaa20f)

One unified canvas. Pick a model card to set your generation style, or a task chip to jump straight to a workflow. The right panel adapts based on what you're doing.

![AI Image Generator canvas showing two generated skincare product images side by side with a glow highlight on the selected one](/files/703b409424688b1b8134bf84ca2a011efa850adc)

Generated images appear directly in the canvas at full size. You can generate multiple variants side by side and compare before downloading.

![Hover state on a generated image showing action icons: expand, wand/edit, thumbs up, thumbs down, download, and a Create video button at the bottom](/files/0dde36d8088ccbf5618a9e054058424c626879fc)

Hover any image to reveal the action row: **Expand** (full-screen view), **Edit** (inpaint, upscale, remove background, variation), **Feedback** (thumbs), **Download**, and **Create video** — which hands the image off directly to the Video Generator.

| Was                                                     | Now — and why it changed                                                                                                                 |
| ------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| Create Images → **Text To Image** (separate page)       | Left rail → **Image** — the default view is text-to-image                                                                                |
| Create Images → **Image To Image** (separate page)      | Same page → click **＋ From a reference image** chip — upload a reference and the controls adapt automatically                            |
| Create Images → **Image Editing** (separate page)       | Same canvas → hover a generated image → click the **Edit (wand)** icon                                                                   |
| Model dropdown in the left panel                        | **Model cards** in the canvas: Flux 1.1 Pro, Gemini · Nano Banana, GPT Image, Ideogram — with subtitles explaining each model's strength |
| Select Preset, Private Session toggle                   | Still available under **Advanced** in the right panel                                                                                    |
| Magic Prompt, Keywords, Negative Keywords               | **Magic Prompt** button in the prompt bar; Negative prompt under **Advanced**                                                            |
| No. of Generations 1–4, Quality 0.5K–4K, Dimension grid | **Number of images** 1–4, **Quality**, **Aspect ratio** grid — same options, right panel                                                 |


# Video Generation

Video generation was tucked inside the **Audio & Video** submenu, alongside Text To Speech and Text To Music. This grouping made sense architecturally, but it meant video — one of the most used tools — required two clicks from the top of the nav.

Video now has its own first-class spot in the left rail. The workspace follows the same unified approach as Image: model cards let you pick the generator (each with different strengths), task chips set the mode (Text→Video, Animate an image, Start & end frame), and settings live in the right panel.

### Before

![Old Video Generation page under Audio & Video submenu showing Model/Veo 3.1 Fast, Reference Images 0/3 upload area, prompt box, Magic Prompt, Keywords in left panel. Right panel shows No. of Generations 1-4, Duration 4/6/8s, Quality 720P/1080P/4K, Dimension 16:9/9:16.](/files/2c16df1ef90ca791f576dc3d606238a27b0d6c87)

Video lived inside a sub-menu, two clicks deep. The interface was split between a left settings panel and a black canvas with collapse/expand arrows.

### After

![New AI Video Generator showing model cards (Seedance 2.0, Kling V3 4K, Veo 3.1, Runway Gen 4.5) with capability subtitles, task chips (Text→Video, Animate an image, Start & end frame), + Reference files button, and right panel with Duration 4s/6s/8s, Resolution 720P/1080P/4K, Aspect ratio 16:9/9:16, Advanced.](/files/15ebbcc0e6765b1788bd80cc5cb8aecf2af28eb8)

Video is now a top-level destination. Model cards show each model's specialty (Refs · audio, Elements, Start · end frame, Cinematic) so you can make an informed choice. The prompt lives at the bottom — describe what you want and hit Generate.

![Video Generator canvas showing a generated white sneaker product video with playback controls (play, mute), expand button, thumbs up/down, and download icon. Below the video is the prompt text.](/files/719f72558c34c5e6f10d16cdf260ab99d38de243)

Generated videos render inline with a full playback control bar. You can rate the output, download it, or immediately feed it into another prompt to iterate — all without leaving the page.

| Was                                                                           | Now — and why it changed                                                                      |
| ----------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |
| Audio & Video → **Video Generation** (2 clicks deep)                          | Left rail → **Video** (1 click from anywhere)                                                 |
| Model dropdown: `Veo 3.1 Fast`                                                | Model cards: **Seedance 2.0, Kling V3 4K, Veo 3.1, Runway Gen 4.5** — with labelled strengths |
| Reference Images 0/3 upload area                                              | **＋ Reference files** button; task chips set the mode automatically                           |
| No. of Generations 1–4, Duration 4/6/8s, Quality 720P–4K, Dimension 16:9/9:16 | All identical in the right panel — same values, same options                                  |
| Magic Prompt, Keywords                                                        | Magic Prompt in prompt bar; Keywords under Advanced                                           |


# Speech Generation

Text To Speech was a standalone page. Text To Music was another standalone page. Multi-speaker dialogue (which required a paid plan) was a tab inside Text To Speech. Three separate destinations for what is really one continuous surface: *turning text into sound*.

The new **AI Audio Studio** unifies all of this. When you arrive, you choose a *mode card*: Single speaker, Multi-speaker, or Music. Everything else — the prompt box, the right panel controls, the model row — reshapes itself around your choice. One page, three modes, zero context-switching.

### Before

![Old Text To Speech page showing Generated Audios list on the left with past entries and voice labels, and Single-Speaker/Multi-Speaker tabs. Right area shows Voice + Accent selectors, Style Prompt field, Your Script textarea, and Get Started With chips.](/files/731d49bdc4352df63096908bd3a23835e768465a)

The old Text To Speech page had your generation history on the left and controls on the right — a fine layout, but it shared no visual connection with Text To Music, which was an entirely separate page.

### After

![New AI Audio Studio showing three mode cards at top: Single speaker (blue mic icon), Multi-speaker (green speech bubble icon), Music (orange note icon). Below the cards is a model row with Flash TTS, Pro TTS, Flash v2.5 ElevenLabs. Right panel shows Model dropdown, Voice, Accent, Style Prompt fields. Prompt box at bottom reads 'Write what you want the voice to say'.](/files/4f564bb5591b5cae0dc5e7efb892b9b1342cc5a8)

The new Audio Studio presents the three modes as equal choices. Pick a mode and everything — the prompt, the model row, the right panel — adapts to it. Single speaker and multi-speaker share the same voice controls; Music replaces them with track settings.

### Multi-speaker

Multi-speaker mode lets you script a dialogue with multiple voices — ideal for podcast-style content, interview simulations, explainer videos, or any content that needs more than one narrator. You assign a distinct voice and gender to each speaker role, then write the script with speaker labels.

#### Multi-speaker mode selected

![Audio Studio with Multi-speaker mode card highlighted. The prompt box now shows a sample multi-speaker dialogue with Speaker 1 and Speaker 2 lines. Generate button at bottom right.](/files/606d03bc8f33ba5175f2d1f759e7792b81a978a8)

Selecting Multi-speaker updates the prompt box with a dialogue template, so you can see the expected format immediately. The right panel switches to show per-speaker voice assignment controls.

#### Generated output

![Multi-speaker output showing waveform audio player with speaker labels (Speaker 1: Zephyr, Speaker 2: Puck) visible above the waveform. Play, download, and share controls below.](/files/36f8aad01696abeabfa8a4eb7be99dadfe1dded3)

The generated audio file shows each speaker labelled with their voice name above the waveform. You can listen back, download the full clip, or regenerate individual segments if one voice needs adjusting.

| Was                                                  | Now — and why it changed                                                                                      |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| Audio & Video → **Text To Speech** (separate page)   | Left rail → **Audio** → **Single speaker** mode (default on arrival)                                          |
| Single-Speaker tab / Multi-Speaker (Pro) tab         | Three mode cards: **Single speaker · Multi-speaker · Music** — all visible, all equal                         |
| Model dropdown: `Flash TTS (faster)`                 | Model row: **Flash TTS (Gemini), Pro TTS (Gemini), Flash v2.5 (ElevenLabs)** — same models, clearer labelling |
| Voice + Accent selectors, Style Prompt, Your Script  | Same controls in the right panel and prompt area — nothing removed                                            |
| Get Started With chips (Narration, Social media ad…) | Same starter chips below the prompt box                                                                       |
| Generated Audios list in left sidebar                | Sidebar **Recent** list + **History** button top-right of the Audio Studio                                    |


# Music Generation

Text To Music was one of the three tools inside the Audio & Video submenu. It had its own dedicated page with a model selector, prompt box, Instrumental toggle, Negative Prompt, Magic Prompt, and Seed. All of that still exists — it's just reached differently now.

Music is the third mode card in the Audio Studio. Click it and the page reconfigures: the prompt box asks for style description, the right panel switches to music-specific controls (Lyria model selection, Instrumental only toggle, Lyrics written for you indicator, Advanced with Negative prompt and Seed). The full-length song generator (Lyria 3 Pro) and the 30-second clip generator (Lyria 3 Clip) are both there.

### Before

![Old Text To Music page showing Model/Lyria 3 Clip dropdown, Prompt field with placeholder 'e.g. Ambient music with synthesizers', 0/5000 character count, Instrumental only checkbox, Negative Prompt field, Magic Prompt section, Seed field, and Generate button with 11 credit cost.](/files/d03001e96d41270c246ed0702e33b29ba0b650f5)

Text To Music was its own nav item and page. Good for discoverability, but disconnected from Speech — even though both are audio generation tools that share models and controls.

### After

![New Audio Studio with Music mode card highlighted (orange icon). Model row shows Lyria 3 Clip and Lyria 3 Pro with descriptions. Right panel shows Lyria 3 Clip selected, Instrumental only toggle, Lyrics written for you label, Advanced section with Negative prompt and Seed. Prompt box reads 'Describe the style: genre, mood, instruments, tempo...'](/files/968b7548cbc609394d31d61e4fa583d6fdae5ccb)

Music as the third mode in Audio Studio. The model descriptions help you choose — Lyria 3 Clip for quick 30-second clips, Lyria 3 Pro for full-length songs with vocals and arrangements. All other controls are identical to before.

| Was                                                                        | Now — and why it changed                                                                                 |
| -------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| Audio & Video → **Text To Music** (separate nav item)                      | Left rail → **Audio** → click **Music** mode card                                                        |
| Model dropdown with `Lyria 3 Clip`                                         | Model row: **Lyria 3 Clip** (30s clips) and **Lyria 3 Pro** (full songs) — same models, described inline |
| Instrumental only checkbox, Negative Prompt, Magic Prompt, Seed (optional) | All present: Instrumental toggle in right panel; Negative prompt + Seed under **Advanced**               |
| Minimum 2–3 words prompt requirement                                       | Same minimum — shown as a hint below the prompt box                                                      |
| Generate button with credit cost shown                                     | Generate button with credit cost shown — unchanged                                                       |


# Account, Credits and Workspaces

Credits and workspace management were buried in the sidebar. Credits & Billing was a nav item you'd visit occasionally; the org switcher was inside the top-right profile dropdown. Neither surfaced the information you needed most — how many credits do I have right now? which workspace am I in? — without clicking away from what you were doing.

In the new design, both are always visible at the top of the left sidebar. The workspace block shows your org name, workspace name, tier, and current credit balance. You never have to leave your current page to check credits or switch orgs.

### Before

![Old Credits & Billing page showing Organization Credits & Billing heading, Qolaba Organization Monthly Credits section with 5939 credits remaining and 49 members, and a Credit Utilization table with Date & Time, Type, Quantity columns listing recent usage entries.](/files/a527610ba73e0ff12b187ac237fcab05ac0d2ea8)

Credits were a page you visited — you couldn't see your balance at a glance. The utilization table was useful for auditing spend, but required navigating to it.

### After

![New Home showing sidebar with Qolaba Team · Star workspace block at top showing ORGANIZATION Tier 2 and Credits 65.5k in real time](/files/f3c9b81c36a39b7990d89c8c521e1636e2100b18)

Your credit balance and workspace are always visible in the sidebar. No navigation required — the information is ambient. The full Credits & Billing detail page (utilization table, top-up, invoices) is one click away from the workspace block.

| Was                                                     | Now — and why it changed                                                                                 |
| ------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
| Left nav → Credits & Billing (full-page, navigate away) | Credit balance always visible in the **sidebar workspace block**; full billing page via the block's menu |
| Org switcher inside the top-right profile dropdown      | **Workspace block at the top of the sidebar** — shows org, workspace, tier, and credits at a glance      |
| Refer & Earn as a sidebar nav item                      | Accessible from the profile/account menu                                                                 |
| Credit Utilization table (Date, Type, Quantity)         | Same table — accessible from the workspace block menu → Billing                                          |


# What remains unchanged ?

The redesign was about navigation and layout, not about changing what Qolaba can do. Here is everything that remained unchanged:

<table><thead><tr><th width="299.12890625">Area</th><th>What stayed the same</th></tr></thead><tbody><tr><td><strong>All AI models</strong></td><td>Every LLM (Gemini, Claude, GPT), image model (Flux, Nano Banana, Ideogram), video model (Seedance, Kling, Veo, Runway), TTS model (Flash TTS, Pro TTS, ElevenLabs), and music model (Lyria) — same providers, same tiers, same access</td></tr><tr><td><strong>Chat history</strong></td><td>Every conversation is intact. Same filters (All, Images, Audio, Video, Favourites), same per-item actions (3-dot menu), same share links</td></tr><tr><td><strong>Agents</strong></td><td>All pre-built and custom agents are unchanged. Agent system prompts, toolkit permissions, sharing settings, and conversation starters are exactly as you left them</td></tr><tr><td><strong>Knowledge bases</strong></td><td>All KBs migrated with their sources, embeddings, names, and sharing settings intact. Ready status means they are immediately usable</td></tr><tr><td><strong>Saved prompts &#x26; files</strong></td><td>All prompts and uploaded files are in Context — same content, new location</td></tr><tr><td><strong>Credits &#x26; plan</strong></td><td>Credit balance, plan tier, monthly allocation, top-up history, invoices, and billing settings are all unchanged</td></tr><tr><td><strong>Organisations &#x26; workspaces</strong></td><td>Members, roles (Owner / Admin / Member), workspace credit allocation, team invitations, and SSO settings are unchanged</td></tr><tr><td><strong>API &#x26; integrations</strong></td><td>All API endpoints, authentication keys, request/response shapes, and OpenAI-compatible routes work exactly as before — no code changes required</td></tr><tr><td><strong>Keyboard shortcuts</strong></td><td>All shortcuts are unchanged. Press <code>?</code> anywhere in the app to see the full shortcut reference panel</td></tr><tr><td><strong>Chat Branching</strong></td><td>Same tree view, same shortcut <code>⌘⇧B</code>, same merge and collapse controls</td></tr><tr><td><strong>Magic Prompt</strong></td><td>Available in Image and Video prompt bars — same one-click prompt enhancement, same credit cost</td></tr><tr><td><strong>Multi-speaker audio</strong></td><td>Same voice-per-role assignment, same output format, same model options — now under Audio → Multi-speaker</td></tr><tr><td><strong>Private sessions</strong></td><td>Available under Advanced in Image generation — same behaviour, no history saved</td></tr></tbody></table>


# FAQs

<details>

<summary>Where's the Text To Music page?</summary>

It's now the **Music** mode inside the Audio Studio. Click **Audio** in the left rail, then click the **Music** mode card. You'll find the same Lyria 3 Clip and Lyria 3 Pro models, the same Instrumental toggle, the same Negative Prompt and Seed fields under Advanced — nothing was removed.

</details>

<details>

<summary>Where did my Knowledge Bases go?</summary>

They're in **Context → Knowledge bases**. Click Context in the left rail and select the Knowledge bases tab. Every KB you had is there, with all its sources, embeddings, and sharing settings intact. The status badge should read "Ready" — meaning it's immediately available to attach to any chat.

</details>

<details>

<summary>Where are my saved prompts and uploaded files?</summary>

Both are in **Context**. Click Context in the left rail, then switch between the **Prompts** tab and the **Files** tab. Your data hasn't moved or changed — just the address has.

</details>

<details>

<summary>Where's the New Chat button?</summary>

It was removed because starting a new chat is now even faster without it. Three options: (1) type directly in the Home prompt box, (2) click **＋** next to Chat in the left rail, or (3) press `⌘⇧O` on Mac / `Ctrl⇧O` on Windows.

</details>

<details>

<summary>Where's the old Image To Image mode?</summary>

Go to **Image** in the left rail and click the **＋ From a reference image** task chip. Upload your reference image(s) and the controls will adapt — the generation settings in the right panel stay the same.

</details>

<details>

<summary>Where are Inpaint, Upscale, Remove Background, and Variation?</summary>

These are still in Image Generation — they just open differently now. Generate an image first, then hover over it. A row of icons appears: the **wand (Edit)** icon opens the editing tools. Inpaint, Upscale, Remove Background, and Variation are all there.

</details>

<details>

<summary>Did my credits or subscription plan change?</summary>

No. Your credit balance, plan tier, and monthly allocation are identical. The only change is where you *see* them — the credit balance is now always visible in the sidebar workspace block, and the full billing page (utilization table, invoices, top-ups) is one click from the workspace block menu.

</details>

<details>

<summary>Do my existing share links still work?</summary>

Yes. Every share link generated before the redesign still resolves to the correct content. Links to specific chats, generated images, agents, and knowledge bases all continue to work.

</details>

<details>

<summary>I have API integrations. Do I need to update anything?</summary>

No changes required. The redesign was purely visual and navigational — all API endpoints, authentication keys, request schemas, response formats, and OpenAI-compatible routes are identical. Your integrations will continue to work without any code changes.

</details>

<details>

<summary>How do I get back to the old Community / Qommunity page?</summary>

The Community is still accessible — look in the **Discover** section, which can be reached from the profile / account menu.

</details>

## Still can't find something?

This guide covers every major change in the redesign, but if you're looking for something specific and it's not listed here, we want to know. The fastest way to get help is to describe what you were trying to do and where you used to find it — that context helps us point you directly to the new location (and improve this guide for everyone else).

* **Email support:** <support@qolaba.io>
* **Full docs:** [docs.qolaba.ai](https://docs.qolaba.ai/)

{% hint style="info" %}
**Reporting a missing feature?** Please include: what you were trying to do, which page or workflow you used before, and what result you expected. Screenshots help a lot. This makes it much faster for us to confirm whether it's a navigation change covered in this guide, a known issue, or something we need to investigate.
{% endhint %}


# Video Series

Create multi-scene AI videos from a single idea, scene-by-scene prompts, or a ready-made template. Video Series handles the planning, generation, and assembly, you control the story. This guide walks you through every workflow, screen by screen, with annotated visuals so you can start producing polished video content in minutes.

### Overview

[**Video Series**](https://www.qolaba.ai/app/video-series) is Qolaba's multi-scene video creation tool. While the standard **Video** generator produces a single clip from a single prompt, Video Series lets you build a *sequence* of clips, each with its own scene description, voiceover direction, and camera cues, and then stitch them together into one continuous video with transitions.

The tool lives in the left navigation rail as its own top-level item, right below Audio and above Context. It is available to all users and carries a FREE badge while in beta.

There are three ways to start a series, each designed for a different level of creative control:

| Starting Mode               | How it works                                                                                                                                                                                               |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **“I have an idea”**        | Describe your concept in plain language. Qolaba's AI plans every scene for you, environment, camera, voiceover, audio cues, and generates a complete scene breakdown automatically.                        |
| **“Scene by scene”**        | Start with a blank canvas. You write each scene prompt yourself, upload reference images for character and product consistency, and maintain full creative control over every detail.                      |
| **“Start from a template”** | Choose a pre-built template (Product review UGC, Activewear try-on, Cinematic Beverage Commercial, etc.), drop in your own product image, and generate immediately, the scene prompts are already written. |


# Getting Started

Navigate to **Video Series** in the left rail. The landing screen presents a clear question: *“What are we creating today?”* with two mode cards at the top and a prompt area below. Beneath the prompt area, a horizontal carousel of template cards gives you a third, shortcut-based starting point.

![Video Series home screen showing 'What are we creating today?' heading, I have an idea and Scene by scene mode cards, prompt area with sample text about a thylacine, starter chips, configuration options, and template carousel below](/files/0da400da8796e75948bfdfac41e0633e6df2165a)

The Video Series landing page. The two mode cards (**I have an idea** and **Scene by scene**) sit at the top. Below the prompt box, quick-start chips offer curated story ideas. Configuration pills (model, orientation, duration) and the **Generate scenes** button line the bottom. Scroll down to find the template gallery.

#### Three Starting Modes

The two mode cards at the top toggle between AI-planned and manual creation:

* **I have an idea** (selected by default): “We plan every scene for you.” Type your concept into the prompt box and hit **Generate scenes**. Qolaba's AI breaks your idea into individual scenes with detailed prompts, camera directions, voiceover scripts, and audio cues.
* **Scene by scene**: “Blank scenes, you fill them in.” Clicking this takes you straight to a blank scene editor where you write each scene's prompt manually and upload reference images for consistency.
* Below the main creation area, a **Start from a template** carousel offers pre-built series. Each template card shows a preview thumbnail, a clip count badge (e.g., “2 clips” or “1 clip”), and a label like “Product review UGC” or “Cinematic Beverage Commercial.”

#### Configuration Options

Before generating, you can configure the series using the pills below the prompt:

| Setting         | Options                                                                         |
| --------------- | ------------------------------------------------------------------------------- |
| **Model**       | `Omni video model`: Qolaba's default video generation model (Gemini Omni Flash) |
| **Orientation** | `Portrait` or `Landscape`: sets the aspect ratio for all clips in the series    |
| **Duration**    | `10s` per clip (configurable via dropdown)                                      |

You can also use the quick-start chips, **Under the Sea**, **Whale Fall**, **Buried History**, **Storm Season**, **Last of its Kind**, to pre-fill the prompt with a curated story idea and jump straight in.

### 3. The Three-Step Pipeline

Regardless of how you start, every Video Series follows the same three-step pipeline. A progress stepper at the top of the workspace tracks where you are:

1. Scenes → 2. Clips → 3. Assemble

| Step            | What happens                                                                                                                                                                                                                               |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **1. Scenes**   | Define the story. Each scene gets a detailed text prompt describing the environment, action, camera work, voiceover, and audio. You can add reference images that carry through every scene for character and product consistency.         |
| **2. Clips**    | Generate the video. Each scene is rendered into an individual video clip. Clips generate one at a time, and you can regenerate any clip independently without affecting the others. A progress badge (e.g., “4/4 done”) tracks completion. |
| **3. Assemble** | Combine into final video. Arrange your clips on a timeline, choose transitions between them (Hard Cut, Dissolve, Fade, Wipe, Slide), preview the result, and combine into a single downloadable MP4.                                       |

You can move backward at any point, the **Edit scenes** and **Back to clips** buttons let you revise earlier stages without losing work. Green checkmarks on completed steps confirm your progress.


# 'I Have an Idea" Flow

This is the fastest path from concept to finished video. You describe what you want in natural language, and Qolaba's AI handles the creative planning, breaking your idea into scenes, writing detailed prompts for each one, and structuring camera directions and voiceover cues automatically.

#### Step 1: Enter your idea

With the **I have an idea** mode card selected, type your concept into the prompt box. Be as descriptive or as brief as you like, the AI will expand on it. Set your preferred orientation and duration, then click the orange **Generate scenes** button (or press `Enter`).

![Prompt box filled with a detailed story concept about a thylacine at Hobart Zoo, with the 'Last of its Kind' chip highlighted in orange and configuration pills showing Omni video model, Portrait, and 10s](/files/0da400da8796e75948bfdfac41e0633e6df2165a)

A rich story concept typed into the prompt box. The “Last of its Kind” chip is highlighted, showing it was used as a starter. The prompt describes a specific historical scenario, the AI will decompose this into individual scenes with camera directions and audio cues.

#### Step 2: AI plans your scenes

After clicking Generate scenes, a loading state appears: *“Planning your scenes… Structuring your story, one scene at a time.”* The AI analyses your idea and creates a structured scene breakdown.

Loading

![Loading screen showing 'Planning your scenes...' with a spinner and subtitle 'Structuring your story, one scene at a time.'](/files/d815606a152ab249085c57af9fd4a6a31ad92e02)

The AI is decomposing your idea into individual scenes. This typically takes a few seconds.

Result

![Generated scene breakdown showing Scene 1 and Scene 2 with detailed prompts including environment descriptions, camera directions, voiceover cues, and audio notes](/files/ab785356f00b8ad9c4b1bb3471e1eac9ad09fc37)

The AI has generated two scenes, each with a rich prompt covering environment, character action, camera movement, voiceover delivery, and ambient audio. A **Carried through every scene** section at the top lets you add a reference image for visual consistency.

Each generated scene prompt is fully editable. You can modify the text, delete scenes with the × button, or add new scenes with **+ Add a scene**. The footer shows the clip count and confirms the free beta pricing.

#### Step 3: Generate clips

Once you are satisfied with the scene breakdown, click the orange **Generate** button. Qolaba renders each scene into a video clip. The pipeline stepper advances to Step 2 (Clips), and you see your generated clips side by side with status badges.

![Clips view showing two generated video clips of a thylacine in black and white, each with a green 'Done' badge, Regenerate button, and scene description preview. Pipeline stepper shows Scenes (checked) and Clips (active)](/files/f19d4d6ab1295f741c8b21c31da1b7fca23db3df)

Both clips have generated successfully (“2/2 done”). Each clip card shows a video thumbnail with playback controls (audio, download, expand), a green **Done** badge, a **Regenerate** button, and a truncated scene description. The breadcrumb reads “Series › Last of it's kind”, the series was automatically named from your idea. The header shows metadata: Portrait, 10s, 2 scenes, Omni Flash.

#### Step 4: Assemble the video

When you are happy with your clips, click **Assemble video →** to move to the final stage. The Assemble view shows your clips on a visual timeline with transition indicators between them. Set transitions, then click **Combine into final video**.

![Assemble stage showing pipeline stepper with Scenes and Clips checked, Assemble active. Timeline with two clip thumbnails numbered 1 and 2, Hard Cut transition indicator between them, Back to clips and Combine into final video buttons, and preview area](/files/727c601292f0e890386a2ae488ea7b35a14d7588)

The Assemble stage for the “I have an idea” flow showing two clips in the timeline with a “Hard Cut” transition between them. Click the transition indicator to change it, or click **Combine into final video** to render. Use **Back to clips** if you need to regenerate any clip before combining.


# "Scene by Scene" Flow

This flow gives you full manual control. Instead of describing a single idea and letting the AI plan the scenes, you start with a blank canvas and write each scene prompt yourself. This is ideal when you have a specific shot list, storyboard, or creative brief that you want to follow precisely.

#### Selecting Scene by Scene mode

On the Video Series home screen, click the **Scene by scene** mode card on the right. The subtitle reads: “Blank scenes, you fill them in.” Set your orientation and duration, then click the orange **Start blank** button.

![Video Series home with 'Scene by scene' card selected, showing 'Blank scenes, you fill them in' subtitle, configuration pills for Omni video model, Portrait, Landscape, 10s, and orange 'Start blank' button](/files/8dde30fdcfecd95bf5af1a8e40cb334ce2b6a851)

The **Scene by scene** card is selected (outlined in orange). The prompt area is replaced by configuration pills and a **Start blank** button. Templates are still visible below for quick access.

#### The blank canvas

You land on the Scenes stage (Step 1) with a single empty scene card. The interface has two main sections:

* **Carried through every scene**: A reference image area at the top where you upload images of your character, product, or subject. These images maintain visual consistency across all generated clips.
* **Your scenes**: Below, numbered scene cards with text areas where you write each scene's prompt. A **+ Add a scene** button lets you add more.

![Blank scene editor showing 'Scenes' heading, a banner reading 'Enjoy the free beta with unlimited clips!', the 'Carried through every scene' reference image area with an '+ Add image' placeholder, one empty Scene 1 card with 'Describe this scene...' placeholder, and footer showing '0 clips - Free' with '+ Add a scene' and 'Generate Free' buttons](/files/97905d9047a006eb7ac05bad5e70b3e92667fd4a)

A fresh Scene by scene canvas. The reference image area (“Carried through every scene”) is empty, ready for you to upload a character or product photo. Scene 1 has a blank text area. The footer confirms 0 clips and free generation.

#### Adding reference images

Click the **+ Add image** placeholder in the “Carried through every scene” section. A modal opens where you can drag-and-drop an image (up to 20 MB) or browse your files. You can also choose from your generation history, filtered by All, Generated, or Uploaded.

![Add reference image modal showing drag-and-drop area with 'Maximum size: 20MB', a 'Choose from history' section with All/Generated/Uploaded tabs, and a grid of previously generated and uploaded images](/files/09fc96c5f6f9ef778cdce50b6bea3b926a128442)

The reference image picker. Upload a new file or choose from your existing image history. The tabs let you filter between AI-generated images and manual uploads.

#### Writing scene prompts with reference images

After adding reference images, each one gets a name label (editable) and appears in the “Carried through every scene” section. You can add multiple reference images, for example, a character photo labelled “Sam” and a product photo labelled “Classic watch.” These references are automatically incorporated into every scene's generation to maintain visual consistency.

![Scene editor with two reference images uploaded: a portrait photo labelled 'Sam' and a watch photo labelled 'Classic watch', plus an '+ Add image' slot. Below, four scene cards with detailed prompts describing Sam discovering a vintage watch, nostalgic flashback, returning to present, and walking down a street](/files/440c1c74ff1cef89ab70d6667467ba6894d725d7)

A fully populated scene editor. Two reference images (“Sam” and “Classic watch”) are set to carry through every scene. Four scenes describe a narrative arc about discovering a vintage watch. Each scene prompt includes cinematic direction: camera movement, lighting, depth of field, and film grain.

#### Editing reference image descriptions

Qolaba automatically generates a text description of each uploaded reference image using AI vision. You can edit this description to refine what the model focuses on, skin tone, clothing, distinguishing features, product details. Click the pencil icon next to the image label to open the description editor.

#### Generating and reviewing clips

Once your scenes are written and reference images are set, click the orange **Generate** button. Qolaba renders each scene into an individual video clip. The pipeline stepper advances to Step 2 (Clips), and you see all generated clips in a grid with status badges.

![Clips stage showing four generated video clips in a grid, each with a green Done badge, Regenerate Free button, and scene description preview. Scene 3 shows Takes 1 and 2.](/files/0e4a5c715daf6bbb9a13b3261fbaa553c42c6f29)

All four clips have generated (4/4 done). Each clip card shows a video thumbnail with playback controls, a green **Done** badge, and a **Regenerate** button. Scene 3 already has two Takes (indicated by the numbered circles), showing it was regenerated for a better result. Click **Assemble video →** to proceed to the final step.

#### Regenerating in Scene by Scene

The regeneration workflow is identical to the “I have an idea” flow. Click **Regenerate** on any clip to open the Edit Scene modal, where you can refine the prompt and create a new take.

![Edit Scene 1 modal in Scene by Scene flow with clip preview, editable prompt about Sam discovering a watch, info banner, and Regenerate button](/files/5ba7c59f4d54a4b0fcc8a818719a9b16740887eb)

The Edit Scene modal for Scene 1 in the Scene by Scene flow. The prompt contains the full cinematic description you wrote manually. You can refine it, adjust camera angles, change the lighting mood, add detail, then click **Regenerate**. Your original clip is preserved as a separate take.

#### Setting transitions in Scene by Scene

On the Assemble stage, click any gap between clips in the timeline to open the Transition Preview. The same seven transition types are available across all flows.

![Transition Preview modal in Scene by Scene flow showing live preview, seven transition type buttons with Hard Cut selected, and Apply to all scenes button](/files/c4e493daf5e68b72d37f66038026de65f974e485)

The Transition Preview modal in the Scene by Scene flow. A live preview shows the actual footage from your clips with the selected transition applied. Choose a type, preview it, and click **Apply to all scenes** to set the same transition everywhere, or set each gap individually.


# Start with a Template

Templates are the fastest way to create a professional video series. Each template is a pre-built series with scene prompts already written by Qolaba's team. You just drop in your own product or character image, and the template adapts the prompts to match. Templates are ideal for commercial content, social media ads, product showcases, and UGC-style videos.

#### The template gallery

Scroll down on the Video Series home screen (or click a template card directly) to browse available templates. Each template card shows a video preview thumbnail, the number of clips it contains, and a descriptive label.

![Video Series page scrolled to show the template carousel with four visible cards: Product review UGC (2 clips), Activewear try-on UGC (2 clips), Cinematic Beverage Commercial (1 clip), Street Poster Campaign (1 clip). Below, a 'Recent series' section shows existing series with status filters (All, In progress, Ready, Drafts) and sort options](/files/ef982518468b057a86295247448d561f9c4b05e4)

The template carousel and the Recent series dashboard below it. Available templates include **Product review UGC**, **Activewear try-on UGC**, **Cinematic Beverage Commercial**, and **Street Poster Campaign**. The Recent series section shows your existing work with status filters and sort controls.

#### Selecting and customising a template

Click a template card to open its setup modal. The modal shows a video preview of the template on the left, and a configuration panel on the right where you upload your product image, choose orientation, and generate.

Step 1: Select & Upload

![Cinematic Beverage Commercial template modal showing a water-drops video preview on the left, and on the right: YOUR IMAGES area with empty '+ Add image' placeholder, 'Product \*' label indicating required, ORIENTATION toggle between Portrait and Landscape, and 'Generate video' button (greyed out)](/files/7d1b41f6fb8d4b7c437d4e3b2b02a9d457413338)

The template modal for “Cinematic Beverage Commercial.” The subtitle reads: “Single high-end commercial clip: macro drops, splash finale.” Upload your product image in the **YOUR IMAGES** area, the field is marked as required (\*\*Product \*\*\*).

Step 2: Generate

![Same template modal but now with a Crush soda can image uploaded and marked with a green checkmark and 'Described' badge. The Generate video button is now orange/active. A note reads 'Credits are only spent when you generate clips'](/files/1d944f42f52ef5c9952fec32d3a4954c96f04674)

A product image (Crush soda can) has been uploaded and automatically described by AI (“Described” badge). The **Generate video** button is now active. Click it to proceed.

#### Template scene review

After clicking Generate video, you land on the Scenes stage. The template has pre-filled the scene prompts with professional-grade directions. Your uploaded product image appears in the “Carried through every scene” section. You can review, edit, or add scenes before generating clips.

![Scenes stage showing the Crush soda can as a reference image labelled 'Product' in the 'Carried through every scene' section. Below, Scene 1 contains a detailed template prompt: 'STYLE: High-end cinematic beverage commercial. Ultra-realistic macro detail, strong contrast lighting...' and 'FORMAT: Portrait orientation (9:16 vertical aspect ratio)...'](/files/3c9a4498dd641d1dfd8cb34604894f52092395ab)

The template has pre-written Scene 1 with professional direction covering style (cinematic, macro detail, commercial colour grading) and format (9:16 portrait, optimised for TikTok, Reels, and Shorts). Your product image is set as the reference. Click **Generate** to render the clip.

#### Generated template clip

The clip generates and appears in the Clips stage, ready for review. You can regenerate, edit scenes, or proceed to assemble.

![Clips stage showing one generated clip of water droplets on a blue surface, with a green 'Done' badge, 'Regenerate Free' button, and a truncated prompt preview. Header shows 'Cinematic Beverage Commercial' with Portrait, 10s, 1 scene, Omni Flash metadata](/files/35e4b11b74b060936e9330a25caef648d09f010e)

The template clip has been generated (1/1 done). The video shows macro water droplets, exactly matching the “macro drops” style described in the template prompt. You can **Regenerate** for a different take or proceed to **Assemble video**.

#### Regenerating a template clip

Even with a pre-written template prompt, the AI may not nail the exact look you want on the first try. Click **Regenerate** to open the Edit Scene modal, where you can fine-tune the template's prompt before generating a new take.

![Edit Scene 1 modal for the Cinematic Beverage Commercial template, showing a clip preview of the Crush soda can, the full editable template prompt with STYLE and FORMAT sections, info banner about keeping the current clip, and Regenerate button](/files/2bf012330615b0d5354043e62480c1608d4afd35)

Regenerating a template clip. The modal shows the full template prompt, style direction, format specifications, and the AI-described product reference. You can edit any part of the prompt to adjust the visual outcome. The cost badge confirms regeneration is **Free** during beta.


# Adding Reference Images

Reference images are one of Video Series' most important features. They solve the core challenge of multi-scene AI video: keeping characters, products, and visual elements consistent from one clip to the next.

When you upload a reference image in the “Carried through every scene” section, Qolaba's AI analyses it, generates a detailed text description, and injects that description into every scene's generation prompt. This means your character's face, clothing, and build, or your product's shape, colour, and branding, remain visually consistent across all clips.

| Feature                          | Detail                                                                                                                                       |
| -------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| **Multiple references**          | Upload several images, one for a character, one for a product, one for an environment. Each gets its own label and AI-generated description. |
| **Editable labels**              | Name each reference (e.g., “Sam”, “Classic watch”, “Product”). Labels appear in scene prompts for easy identification.                       |
| **AI-generated descriptions**    | The system automatically writes a detailed visual description. You can edit it to emphasise or de-emphasise specific features.               |
| **Set once, applied everywhere** | Reference images are set at the series level, not per scene. Every clip in the series inherits them automatically.                           |
| **Source flexibility**           | Upload from your computer (up to 20 MB), or choose from your Qolaba image generation history (All / Generated / Uploaded tabs).              |


# Regenerating Clips

Not every AI-generated clip will be perfect on the first attempt. Video Series makes iteration painless: you can regenerate any individual clip without affecting the rest of your series. Your current clip is always preserved, the new version is added alongside it as an additional “take.”

#### How to regenerate

On the Clips stage, click the **Regenerate** button on any clip card. A modal opens showing the current clip's video preview and its editable prompt. You can modify the prompt text before regenerating to guide the AI in a different direction.

Edit Prompt

![Edit Scene 1 modal showing the current clip's video preview on the left and an editable prompt field on the right. An info banner reads 'Your current clip is kept. The new version will be added alongside it; other scenes stay untouched.' Footer shows 'Regenerating will cost Free' with Cancel and Regenerate buttons](/files/f1118c03640e27ba0c707de54e9998bc4eef35e2)

The regeneration modal. The info banner confirms: “Your current clip is kept. The new version will be added alongside it; other scenes stay untouched.” The prompt is fully editable, tweak the camera direction, change the lighting, adjust the voiceover, then click **Regenerate**.

Multiple Takes

![Clips view showing Scene 3 with a highlighted card, two take indicators (Take 1 outlined, Take 2 filled orange), a play button overlay on the video, and the Regenerate button. The card is visually expanded compared to other scene cards](/files/e913947521b8483063dec6dafe51977642231e19)

After regeneration, Scene 3 now shows **Takes** indicators, numbered circles that let you switch between versions. Take 1 (outlined) and Take 2 (filled orange) are both available. Click either to preview that version, and the one you select is used in the final assembly.

✓

**Non-destructive workflow.** Regeneration never deletes your previous clip. Every take is preserved, and you choose which one to use when assembling. This lets you experiment freely without risk.


# Final Video and Download

After setting your transitions, click **Combine into final video**. Qolaba stitches all your clips together with the selected transitions and renders a single video. The result appears in a full-width video player with playback controls.

“I have an idea” result

![Final video player showing a black-and-white thylacine video at 0:03/0:20 with play, volume, and fullscreen controls. Below the player: 'Download MP4' button. Above: timeline showing Scene 1 and Scene 2 thumbnails, 'Back to clips' and 'Re-combine' buttons](/files/e026f033284bfbe82786fa5bce1e4b281f18c39a)

A 20-second final video assembled from two 10-second clips in the “I have an idea” flow. The player includes play/pause, seek bar, volume control, and fullscreen. Click **Download MP4** to save the finished video. **Re-combine** lets you change transitions and render again.

“Scene by scene” result

![Final video player showing a cinematic scene of a vintage watch on aged papers at 0:04/0:39 with playback controls. Above: timeline with 4 scene thumbnails, 'Back to clips' and 'Re-combine' buttons. Below: 'Download MP4' button](/files/b0eecd3cc504a686a389d07d5243a94f60a157c4)

A 39-second final video assembled from four 10-second clips in the Scene by scene flow. The longer duration reflects the four-scene narrative. The same playback controls and Download MP4 button are available.

After downloading, you can always return to your series from the **Recent series** dashboard on the Video Series home page. Every series is saved automatically, you can re-enter it, regenerate clips, change transitions, and re-combine at any time.

Template result

![Final video player showing a Crush soda can surrounded by blueberries and leaves, rendered in portrait orientation. 'Download MP4' button below. Timeline shows Scene 1 thumbnail, 'Back to clips' and 'Re-combine' buttons above](/files/1fb160332dda68d3a952698261ae19809f607360)

A portrait-orientation final video from the Cinematic Beverage Commercial template. The AI has generated a lush product shot of the Crush can with berries and foliage, matching the template's “macro drops, splash finale” style direction.

Template assembly

![Assemble stage for the Cinematic Beverage Commercial with one scene in the timeline, 'Back to clips' and 'Re-combine' buttons, and a large video preview showing the Crush can product shot in portrait orientation](/files/eb2fdfa70bb4e3f4f5d0d17984d6eb9dba9d0047)

The Assemble stage for a single-clip template series. Even single-clip series go through the Assemble step, allowing you to add more scenes later if you want to extend the video.


# Managing your Series

The lower half of the Video Series home page is your series dashboard. It shows all your existing series as cards with status indicators, and provides filtering, sorting, and search controls for managing a growing library.

![Video Series dashboard showing template carousel at top, and below it a 'Recent series' section with filter chips (All 3, In progress 1, Ready 0, Drafts 2), sort options (Recent, Name, Progress), and grid/list toggle. Three series cards visible: 'Try-on Videos' with Product Reviews label, 'Cinematic Commercial' (Generating, 40s), and another series (20s)](/files/9aa669710b4cbb77695864ece64381b96726bd60)

The Recent series dashboard showing three series in different states. Filter chips at the top let you view **All**, **In progress**, **Ready**, or **Drafts**. Sort by **Recent**, **Name**, or **Progress**. Toggle between grid and list views with the icons at the right. A search bar lets you find series by name.

| Status                       | Meaning                                                                                                                |
| ---------------------------- | ---------------------------------------------------------------------------------------------------------------------- |
| **Draft**                    | Scenes are defined but no clips have been generated yet. You can return and continue anytime.                          |
| **In progress / Generating** | Clips are currently being generated. The card shows a duration badge (e.g., “40s”) and an orange generating indicator. |
| **Ready**                    | All clips have been generated and the series is ready for assembly and download.                                       |


# Quick Reference

| Action                           | How to do it                                                                           |
| -------------------------------- | -------------------------------------------------------------------------------------- |
| Open Video Series                | Left rail → **Video Series**                                                           |
| Start from an idea               | Select **I have an idea** card → type your concept → **Generate scenes**               |
| Start from scratch               | Select **Scene by scene** card → **Start blank** → write scene prompts                 |
| Start from a template            | Scroll to templates → click a card → upload product image → **Generate video**         |
| Add a reference image            | Click **+ Add image** in “Carried through every scene” → upload or choose from history |
| Edit a reference description     | Click the pencil icon next to the image label → edit the AI-generated text → **Save**  |
| Generate all clips               | Click the orange **Generate** button at the bottom of the Scenes stage                 |
| Regenerate a single clip         | Click **Regenerate** on any clip card → optionally edit the prompt → **Regenerate**    |
| Switch between takes             | Click the numbered **Takes** circles on a regenerated clip card                        |
| Set a transition                 | Click the gap between two clips in the timeline → choose a transition type             |
| Apply transition to all          | In the Transition Preview modal → **Apply to all scenes**                              |
| Combine into final video         | **Assemble video →** button → set transitions → **Combine into final video**           |
| Download                         | Click **Download MP4** below the final video player                                    |
| Re-edit after combining          | **Back to clips** → make changes → **Re-combine**                                      |
| Return to a previous series      | Scroll to **Recent series** on the Video Series home → click any series card           |
| Add more scenes after generating | Click **Edit scenes** → **+ Add a scene** → write prompt → **Generate**                |


# Introduction

### About Qolaba AI

Qolaba AI is a unified AI workspace — a multi-model AI platform designed to consolidate the world's most powerful AI tools into a single, seamless environment. By bringing together 60+ leading models under one roof, Qolaba eliminates the need to manage separate subscriptions for services like ChatGPT, Claude, Gemini, Grok and Perplexity.

The platform is built on a flexible, credit-based pricing system that removes subscription fatigue and vendor lock-in. This pay-for-what-you-use model allows you to stay in control of your budget and scale your AI operations only when you are ready.

{% embed url="<https://youtu.be/XdVv6GJto8s?si=-x_sh3M_931ll-W5>" %}

### What you can Build

Qolaba enables individuals and teams to experiment, build, and deploy AI-powered features across a wide variety of media formats.

<table data-header-hidden><thead><tr><th width="206.859375">Content Type</th><th>Models Available</th></tr></thead><tbody><tr><td>Text &#x26; Chat</td><td>Gemini 2.5 Pro, Gemini 3 Pro, GPT-5.4, GPT- 5.5, Claude Opus 4.6, Claude Sonnet 4.6, DeepSeek V3.2, Grok 4, Grok 4.2, Perplexity Sonar Reasoning, Sonar Deep Research and more</td></tr><tr><td>Images</td><td>Gemini Nano Banana 1, Nano Banana Pro, GPT Image 2, Seedream 4.5, Imagegen 4, Imagegen 4 Fast</td></tr><tr><td>Speech/Audio</td><td>Gemini Flash TTS, Gemini Pro TTS (multilingual — English, Hindi, Tamil, German, Arabic, Mandarin, and more)</td></tr><tr><td>Video</td><td>Veo 3.1, Veo 3.1 Fast, Seedance 2.0, Seedance 2.0 Fast, Happy Horse, Runway Gen 4.5, Kling O3 and V3 Pro, Hailuo 2.3 Pro, Grok Imagine Video</td></tr><tr><td>Custom AI Agents</td><td>No-code agent builder utilizing any LLM combined with a private knowledge base</td></tr></tbody></table>

{% hint style="info" %}
Model availability may vary based on your subscription plan and credit balance. The platform is continuously updated as new models are released.
{% endhint %}

### Who Qolaba Is For

Qolaba is built as an AI platform for teams and individuals across industries — from solo freelancers to enterprise departments.<br>

* **Business Owners** — Streamline operations, generate marketing assets, automate workflows, and leverage AI-driven insights to scale efficiently without hiring large teams.
* **Marketing Agencies** — Produce full multimedia campaigns — from SEO-optimised copy to scroll-stopping visuals and professional voiceovers — using an AI workspace for agencies that eliminates tool-switching.
* **Content Creators** — Create engaging blogs, social media posts, thumbnails, scripts, videos, and AI-generated visuals from one unified AI platform.
* **Freelancers** — Manage client projects end-to-end with AI-powered writing, design, research, and automation tools in a single workspace.
* **Consultants** — Deliver faster research, strategy decks, data analysis, and AI-assisted reports with professional-grade outputs.
* **Developers** — Integrate advanced AI capabilities into applications via a unified API platform, experiment with multiple models, and deploy faster with reduced overhead.
* **Enterprise teams (including DACH / German businesses)** — Work with a platform that provides Data Processing Agreements (DPA / AVV) on Teams and Business plans. See Section 16 for full data handling details.


# Key Concepts - Quick Mental Model

Before you start using Qolaba, it’s helpful to understand the core concepts that power the platform.

<table data-header-hidden><thead><tr><th width="159.21875">Term</th><th>What it means</th></tr></thead><tbody><tr><td>Credits</td><td>The currency used to run AI models. Every action — chat, image, video — costs a fixed number of credits. Credit-based pricing means you pay only for what you use.</td></tr><tr><td>Organization</td><td>Your top-level account. Holds your subscription, credit pool, members, and all workspaces.</td></tr><tr><td>Workspace</td><td>A project-level environment inside an organization. Separate workspaces isolate clients, departments, or projects.</td></tr><tr><td>Models</td><td>The AI engines that power each feature — LLMs for chat, diffusion models for images, etc.</td></tr><tr><td>Agents</td><td>Custom AI personas you build with a system prompt, optional knowledge base, and a chosen model. Agents maintain consistent role behaviour across sessions.</td></tr><tr><td>Toolkit</td><td>A set of capability toggles (web search, file search, image generation, code execution, PII protection, etc.) that extend what the chatbot can do.</td></tr><tr><td>Knowledge Base</td><td>A collection of documents (PDF, DOCX, TXT, images, URLs) the model references before answering in chat.</td></tr></tbody></table>


# Account Setup

How to create your Qolaba account, verify your email, and complete onboarding to start generating with 400 free credits.

Getting started with Qolaba is quick and seamless. This guide will walk you through the process of creating your account, accessing your dashboard, and completing the onboarding steps to help you make the most of the platform from day one.

{% stepper %}
{% step %}

### Sign Up For Free

* Go to [Qolaba.ai](https://www.qolaba.ai/)
* Click **"Signup"**
* Enter your email address and create a password&#x20;

{% hint style="info" %}
You can sign up faster using Google or other social login options.
{% endhint %}
{% endstep %}

{% step %}

### Verify your account

* Check your email inbox for a verification code
* Enter the code to activate your account
  {% endstep %}

{% step %}

### Login and Explore

* Log in at [Qolaba.ai](https://www.qolaba.ai/)&#x20;
* You'll receive **400 free credits upon signup immediately**
  {% endstep %}

{% step %}

### Complete Onboarding

* Once you log in, you’ll land on the **Onboarding Dashboard**
* Follow the guided tasks and challenges to get familiar with the platform
* These steps are designed to help you understand key features and workflows

{% hint style="success" icon="rocket-launch" %}
Completing onboarding helps you navigate the platform more confidently and unlock its full potential faster.
{% endhint %}
{% endstep %}
{% endstepper %}


# Quick Start Guide

This guide walks you through the essential steps to start using Qolaba. Follow along to explore the platform and understand how everything works together.

### **1. Log In & Access Your Dashboard**

* Log in to your [Qolaba](https://www.qolaba.ai/) account
* You’ll land on the **Main Dashboard**

Once you’re in:

* Your **available credits** are visible at the top
* This dashboard is your central place to access all tools and features

***

### **2. Create Your Organization**

* Click your **profile icon (top right)**
* Select **Create Organization**
* In the setup screen:
  * Enter your **organization name**
  * Select the **number of credits** required
  * Choose the **number of people** in your organization
  * Specify your **role** (e.g., Founder, Marketer, Developer, etc.)
* Complete the checkout process
* Once done, you’ll be redirected to your **Organization Dashboard**

{% hint style="warning" icon="lightbulb" %}
You can name your organization based on your work or brand, such as *“Growth Studio”* or *“AI Consulting Co.”*
{% endhint %}

[Learn more about Organizations *→*](/organization-and-workspaces/organizations)

***

### **3. Create a Workspace & Invite Team Members**

Workspaces help you organize your work within an organization and collaborate with your team.

#### **3.1 Create a Workspace**

1. Go to your **Organization Dashboard**
2. Navigate to the **Workspaces** tab
3. Click **Create Workspace**
4. Enter:
   * **Workspace Name**
   * **Description**
5. Click **Create**

Your workspace will be created and ready to use.

{% hint style="warning" icon="lightbulb" %}
You can structure your workspaces based on how you organize your work:

* **By clients/projects:**
  * *“Marketech”*, *“FoodCo”*, *“Project Alpha”*
* **By departments/functions:**
  * *“Marketing”*, *“Product”*, *“Design”*, *“Content”*

This helps you keep work clearly separated and easy to manage.
{% endhint %}

[Learn More about Workspaces *→*](/organization-and-workspaces/workspaces)

***

#### **3.2 Invite Team Members**

1. Go to the **Members** tab
2. Click **Invite Member**
3. Enter one or more **email addresses**
4. Assign a **role**:
   * **Admin** → Full access and control
   * **Member** → Limited access based on permissions
5. Select the **workspace(s)** you want to invite them to
6. Click **Invite**

Learn More about [Inviting Members](/organization-and-workspaces/members-and-role-based-access/inviting-members) and [Role Based Access ](/organization-and-workspaces/members-and-role-based-access)

***

### **4. Explore Chatbot (LLMs)**

The Chatbot lets you interact with different AI models for writing, research, and analysis.

#### **4.1 Get Started**

* Go to [**Chatbot**](https://www.qolaba.ai/chat)
* In the **LLM selector**, choose a model ( [*Available models →*](/model-reference/chatbot-models) *)*
* Enter a prompt like :
  * *“Explain how language models work in simple terms”*
  * *“Break down prompt engineering with real-world examples”*

{% hint style="warning" icon="lightbulb" %}
Use the **microphone option** to speak your prompt instead of typing.
{% endhint %}

***

#### **4.2 Upload a File (Optional)**

* Click the **“+” (Add File)** icon in input box
* Upload a document (PDF, CSV, etc.)
* Ask:
  * *“Summarize this document”*
  * *“Extract key insights from this file”*

{% hint style="info" %}
You can switch between models anytime and compare responses for the same prompt.\
[*Explore chat branching →*](/chatbot/model-selection/chat-branching)
{% endhint %}

[Learn More about Chatbot *→*](/chatbot/introduction)

***

### **5. Create a Knowledge Base**

* Go to Knowledge Bases :book: in left Navigation Panel
* Click "**+ Create New**"&#x20;
* Upload files (PDF, CSV, TXT, DOCx etc.)
* Click **"Add to Knowledge Base"**

{% hint style="info" %}
Knowledge Bases use **RAG (Retrieval-Augmented Generation)** powered by advanced file search to retrieve relevant information and improve response accuracy

[Learn More about Knowledge Bases *→*](/chatbot/adding-context-and-resources/knowledge-bases)
{% endhint %}

***

### **6. Create and Use Your First Agent**

Agents help you create reusable AI assistants with predefined context and behavior.

#### 6.1 Create an Agent

* Go to **Agents Selector** in prompt box
* Click **Create Agent**
* Select a base model
* Add:
  * A **system prompt** (e.g., *“*&#x59;ou are a market research assistant. Analyze data, identify trends, and provide structured insights based on the provided knowledge base.*”*)
  * Your **Knowledge Base**
* Save the agent

#### 6.2 Use the Agent in Chat

* Go to [**Chatbot**](https://www.qolaba.ai/chat)
* Select your created **Agent**
* Start interacting

Try prompts like:

* *“*&#x53;ummarize the key trends from the uploaded report&#x73;*”*
* *“*&#x49;dentify patterns and insights from the data provide&#x64;*”*

[Learn more about Agents *→*](/chatbot/agents)

***

### **7. Generate Your First Image**

* Go to [**Text-to-Image**](https://www.qolaba.ai/ai-image-editor/text-to-image)
* Select a model ( [Available Models *→* ](/model-reference/image-models)*)*
* Choose a **preset/style** (e.g., *Pixel Art, Photographic, Cinematic*)
* Enter a prompt such as:
  * *“A futuristic city at sunset with cinematic lighting”*
* Click **Generate**

{% hint style="warning" icon="lightbulb" %}
Explore the [**Qommunity Wall**](https://www.qolaba.ai/qreative-wall) for prompt ideas, creative usecases and inspiration from other users!
{% endhint %}

[Learn more about Image Generation *→*](/image-generation/introduction)

***

### **8. Create Your First Video**

* Go to [**Text-to-Video**](https://www.qolaba.ai/video-generation)
* Select a model ( [Available Models *→*](/model-reference/video-models) *)*
* Upload a **reference image** to guide the video generation
* Enter a prompt like:
  * *“A cinematic shot of ocean waves at sunrise”*
* Click **Generate**

{% hint style="warning" icon="lightbulb" %}
Use **Magic Prompt** to enhance your prompt and get better results.
{% endhint %}

[Learn more about Video Generation *→*](/video-generation/introduction)

***

### **9. Generate Your First Audio**

1. Go to [**Text-to-Speech**](https://www.qolaba.ai/ai-speech-generator/text-to-speech)
2. Select a **voice**
   * Each voice has a unique tone and style
3. Choose an **accent**
   * Select the dialect based on region
4. Add a **style prompt** *(optional)*
   * Example: *“Speak in a warm and enthusiastic tone”*
5. Enter your **script**
   * Example:\
     *“Welcome to our platform. Discover how AI can help you create faster and smarter.”*
6. Select a **model**:
   * **Flash TTS** → Faster generation
   * **Pro TTS** → Higher quality output
7. Click **Generate**

{% hint style="info" %}
You can switch between **single-speaker** and **multi-speaker** modes.\
*Learn more about* [*Speech Generation Modes →*](/speech-generation/speech-generation-modes)
{% endhint %}

[Learn more about Speech Generation *→*](/speech-generation/introduction)

***

### **10. Check Your Credits & Usage**

* Your **remaining credits** are visible at the **top right corner** of the dashboard
* You can **top up credits** anytime if needed
* To track usage, go to **Dashboard → Overview → Credit Utilization**
  * View **credits used**, **date & time**, and **type of action**

***

### **11. Access Your History**

1. Go to [**History**](https://www.qolaba.ai/history)
2. View your past **Images, Videos and Audio**
3. Use filters to:
   * Filter by **type** (Image, Video, Audio)
   * Filter by **visibility** (Private, Public, Uploaded)
   * *(If in an organization)* Filter by **workspace**
4. Manage items:
   * **Favorite / Unfavorite** using the star ⭐ icon
   * Click the **three-dot menu (⋯)** to:
     * Copy prompt
     * Download&#x20;
     * Copy shareable link
     * Update visibility
     * Delete
     * Share

{% hint style="warning" icon="lightbulb" %}
You can group items and filter by groups to easily organize and manage your history.
{% endhint %}

> **Note :** To access **chat history**, go to [**Chatbot**](https://www.qolaba.ai/chat) — chat history is available there separately.

***

### **You’re Ready to Go 🚀**

By completing this guide, you’ve:

* Set up your **Organization and Workspaces**
* Explored **Chat and multiple AI models**
* Created a **Knowledge Base and Agent**
* Generated content across **Image, Video, and Audio**
* Learned how to manage **Credits and History**


# API Platform

Qolaba is a powerful AI-driven platform for media transformation. Our REST-based API provides a straightforward path to leverage AI models and creative channels. With predictable URLs and JSON-based communication, the Qolaba API is easy to use and navigate.

### Getting Started

1. **Sign Up**: Create an account on the [Qolaba API platform.](https://platform.qolaba.ai/).
2. **Obtain API Key**: Generate an API key to authenticate your requests. You can find the "Create New Key" option on the platform.
3. **Test the APIs**: Qolaba offers free credits for you to explore the platform's capabilities.

<figure><img src="/files/nbcmuB7OeIdsZBdRTSBG" alt=""><figcaption></figcaption></figure>

* **Try the APIs**: After obtaining your API key, you can use the Qolaba API documentation platform to test the available endpoints.
* **Purchase Credits**: For extended usage, you can purchase additional credits on the Qolaba platform.
* **Track Usage**: Each API call will result in a deduction of credits. You can find the details of your usage in the platform's usage section.

<figure><img src="/files/D8PAuQ1w1AoIe6gFoENT" alt=""><figcaption></figcaption></figure>

For detailed information on the available APIs, request parameters, and response schemas, please refer to the subsequent pages of this documentation.


# Text to Image

For more detailed information on the available text-to-image models, please refer to the [Broken mention](broken://pages/qtClo3MhfdW0RsMkVjg1) section.

## Text to Image API

<mark style="color:green;">`POST`</mark> `/getTextToImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name                  | Type   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         |
| --------------------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| app\_id               | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                                                                                                                                                                                                                                        |
| height                | int    | <p>-> The <code>height</code> parameter represents the vertical dimension of an image. </p><p>-> The valid range for the parameter is between 256 and 1536 pixels. </p>                                                                                                                                                                                                                                                                                                                                                             |
| width                 | int    | <p>-> The <code>width</code> parameter represents the horizontal dimension of an image. </p><p>-> The valid range for the parameter is between 256 and 1536 pixels. </p>                                                                                                                                                                                                                                                                                                                                                            |
| num\_inference\_steps | int    | <p>-> The <code>num\_inference\_steps</code> parameter represents the number of denoising iterations to perform during the image generation process. Generally, more iterations can result in higher-quality images, but they also increase the time required for generation.</p><p>-> The valid range for the <code>num\_inference\_steps</code> parameter is between 1 and 50.</p>                                                                                                                                                |
| guidance\_scale       | float  | <p>-> The <code>guidance\_scale</code> parameter determines how closely the generated image adheres to the provided prompt. Higher values result in the model following the prompt more closely, while lower values allow for more creative deviation.</p><p>-> The valid range for the <code>guidance\_scale</code> parameter is between 1 and 30.</p>                                                                                                                                                                             |
| batch                 | int    | <p>-> The <code>batch</code> parameter allows you to specify the number of images to generate at once. </p><p>-> The valid range for this parameter is between 1 and 8.</p>                                                                                                                                                                                                                                                                                                                                                         |
| prompt                | string | -> The `prompt` parameter is the textual input that guides the image generation process. This prompt serves as an artistic compass, shaping the visual output.                                                                                                                                                                                                                                                                                                                                                                      |
| negative\_prompt      | string | -> The `negative_prompt` parameter allows you to specify content that you want the image generation model to avoid or minimize in the output. This can be useful for excluding certain visual elements or styles that you do not want to be present in the generated image.                                                                                                                                                                                                                                                         |
| celery                | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.                                                                                                                                                                                                                             |
| inference\_type       | string | <p>-> The <code>inference\_type</code> parameter allows you to specify the GPU to be used for the image generation task. The supported value is:</p><ul><li><code>a10g</code></li></ul><p>This parameter is only applicable for Qolaba-deployed models, including Turbo Vision, Qolaba Style, Cartoon, Realistic, and Anime Style.</p><p>The different GPU options provide varying levels of performance and capabilities, allowing you to choose the most suitable GPU based on your requirements and the demand for the task.</p> |

**APP ids for different Text to Image styles**

<table><thead><tr><th width="312">App ID</th><th width="174">Model Name</th><th>Height and Width Constraints </th></tr></thead><tbody><tr><td>ap-n2p3fg3gsvbgnYeEEdef</td><td>Image Gen 4</td><td><p>The API supports the following combinations of <code>height</code> and <code>width</code> parameters:</p><ul><li>720x1280</li><li>960x1280</li><li>1024x1024</li><li>1280x960</li><li>1280x720</li></ul><p>Please ensure that the values you provide for <code>height</code> and <code>width</code> match one of these supported combinations. </p></td></tr><tr><td>ap-nOpQr7stuvwxYzABcdef</td><td>SD 3.5 medium</td><td><p>For SD 3.5 Medium,the API supports the following combinations of <code>height</code> and <code>width</code> parameters:</p><ul><li>720x1280</li><li>800x1000</li><li>1024x1024</li><li>1280x720</li><li>1600x676</li></ul><p>Please ensure that the values you provide for <code>height</code> and <code>width</code> match one of these supported combinations.</p></td></tr><tr><td>ap-mNopQ8rstuvwXYZabcde</td><td>SD 3.5</td><td><p>The API supports the following combinations of <code>height</code> and <code>width</code> parameters for SD3.5:</p><ul><li>720x1280</li><li>800x1000</li><li>1024x1024</li><li>1280x720</li><li>1600x676</li></ul><p>Please ensure that the values you provide for <code>height</code> and <code>width</code> match one of these supported combinations.</p></td></tr><tr><td>ap-rStUv6xyzabcdPQRSefg</td><td>SD 3.5 Turbo</td><td><p>The API supports the following combinations of <code>height</code> and <code>width</code> parameters:</p><ul><li>720x1280</li><li>800x1000</li><li>1024x1024</li><li>1280x720</li><li>1600x676</li></ul><p>Please ensure that the values you provide for <code>height</code> and <code>width</code> match one of these supported combinations.</p></td></tr><tr><td>ap-jKlMn5opqzabcXyZtUVw</td><td>Recraft V3</td><td><p>The API supports the following combinations of <code>height</code> and <code>width</code> parameters:</p><ul><li>720x1280</li><li>800x1000</li><li>1024x1024</li><li>1280x720</li><li>1600x676</li></ul><p>Please ensure that the values you provide for <code>height</code> and <code>width</code> match one of these supported combinations.</p></td></tr><tr><td>ap-jXyZa9bcdefghijklmnopq</td><td>Flux Schnell</td><td><p>The API supports the following combinations of <code>height</code> and <code>width</code> parameters:</p><ul><li>720x1280</li><li>800x1000</li><li>1024x1024</li><li>1280x720</li><li>1600x676</li></ul><p>Please ensure that the values you provide for <code>height</code> and <code>width</code> match one of these supported combinations.</p></td></tr><tr><td>ap-fGhKl3mfkdlpqrtsUVWcba</td><td>Flux Dev</td><td><p>For Flux Dev , the API supports the following <code>height</code> and <code>width</code> parameters:</p><ul><li>1056x1440</li><li>1024x1024</li><li>1440x1056</li></ul><p>When providing the <code>height</code> and <code>width</code> values, please ensure that the <code>height</code> and <code>width</code> matches one of the supported combinations in the list above.</p><p></p></td></tr><tr><td>ap-fGhKl3mfkdlpqrsTuvwxYz</td><td>Flux Pro</td><td><p>For Flux Pro, the API supports the following <code>height</code> and <code>width</code> parameters:</p><ul><li>1056x1440</li><li>1024x1024</li><li>1440x1056</li></ul><p>When providing the <code>height</code> and <code>width</code> values, please ensure that the <code>height</code> and <code>width</code> matches one of the supported combinations in the list above.</p></td></tr><tr><td>ap-sdSyd0idsndjnsnsndjsds</td><td>Dalle 3</td><td><p>The API supports the following <code>height</code> and <code>width</code> parameter combinations for DALL-E 3:</p><ul><li>1792x1024</li><li>1024x1024</li><li>1024x1792</li></ul><p>Please ensure that the <code>height</code> and <code>width</code> values you provide match one of these supported combinations.</p></td></tr><tr><td>ap-x7q8hj9kltmNoPqRzabc</td><td>Leonardo</td><td>The height and width parameters must be multiples of 8 pixels.</td></tr><tr><td>ap-hJkLm4nqzxybwvUTSRdca</td><td>Ideogram</td><td>For Ideogram, the API supports the following aspect ratios for the height and width parameters: 16:9 1:1 2:3 3:2 4:3 3:4 9:16 1:3 and 3:1. <br>When providing the height and width values, please ensure that the resulting aspect ratio matches one of the supported ratios in the list above.</td></tr><tr><td>ap-zuzhawbgipcrnxdtefhjbnvhc</td><td>GPT Image</td><td><p></p><p>The API supports the following height and width parameter combinations for GPT Image:</p><ul><li><strong>1024x1536</strong></li><li><strong>1024x1024</strong></li><li><strong>1536x1792</strong><br>Please ensure that the height and width values you provide match one of these supported combinations.</li></ul></td></tr></tbody></table>

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/text-to-image>" %}


# Image Generation

Generate images from text prompts or source images using any supported model. Supports both **synchronous** (wait for result) and **asynchronous** (poll for result) processing modes.

## Image Generation API

**Endpoint:** `POST /api/v1/images/generate`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

### Quick Start

```bash
curl -X POST https://api.platform.qolaba.ai/api/v1/images/generate \
  -H "X-API-Key: your_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "vertex/nano-banana-flash",
    "prompt": "A photorealistic cat sitting on the Eiffel Tower at sunset",
    "aspect_ratio": "16:9"
  }'
```

**Response `200 OK`:**

```json
{
  "task_id": "task_1705312200000_abc123",
  "prompt": "A photorealistic cat sitting on the Eiffel Tower at sunset",
  "images": [
    {
      "url": "https://storage.example.com/image.png",
      "width": 1280,
      "height": 720,
      "content_type": "image/png"
    }
  ],
  "usage": {
    "images_generated": 1,
    "cost_usd": 0.03,
    "cost_credits": 6
  }
}
```

***

### Processing Modes

The `celery` field controls how the request is processed.

| Mode                      | `celery` | HTTP Status    | Description                                                                                                               |
| ------------------------- | -------- | -------------- | ------------------------------------------------------------------------------------------------------------------------- |
| **Synchronous** (default) | `false`  | `200 OK`       | Waits for generation to complete and returns images directly. Best for low-latency use cases.                             |
| **Asynchronous**          | `true`   | `202 Accepted` | Returns a `task_id` immediately. Generation happens in the background. Poll `GET /api/v1/tasks/{task_id}` for the result. |

**When to use async (`celery: true`):**

* Generating large batches
* High-resolution images with longer generation times
* Background jobs / queue-based workflows
* Avoiding client timeout issues

***

### Request Body

```json
{
  "model": "string (required)",
  "prompt": "string",
  "celery": false
}
```

#### Core Fields

| Field    | Type                | Required | Description                                                                          |
| -------- | ------------------- | -------- | ------------------------------------------------------------------------------------ |
| `model`  | `string`            | Yes      | Model ID. See Supported Models.                                                      |
| `prompt` | `string` (max 4000) | Yes\*    | Text description of the image. \*Optional for image-to-image models (e.g. BiRefNet). |
| `celery` | `boolean`           | No       | Processing mode. `false` = sync (default), `true` = async.                           |

#### Media Inputs

| Field              | Type                                               | Description                                                                      |
| ------------------ | -------------------------------------------------- | -------------------------------------------------------------------------------- |
| `image_urls`       | `string[]`                                         | Source image URLs for image-to-image or conversational editing.                  |
| `reference_images` | `{ url: string, description?: string }[]` (max 14) | Reference images for multi-reference consistency (character, product, or style). |
| `mask_url`         | `string (URL)`                                     | Black/white mask indicating the inpainting area (FLUX Inpainting only).          |

#### Size & Resolution

| Field          | Type     | Description                                                                                                                       |
| -------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------- |
| `aspect_ratio` | `string` | Output aspect ratio. See valid values below. Not all models support all ratios.                                                   |
| `quality`      | `string` | Resolution/quality preset. Model-specific values — e.g. `"2K"`, `"4K"` for Vertex; `"low"`, `"medium"`, `"high"` for GPT Image 2. |

**Supported `aspect_ratio` values:**

```
1:1  1:4  1:8  2:3  3:2  3:4  4:1  4:3  4:5  5:4  8:1  9:16  16:9  21:9
```

#### Generation Parameters

| Field                 | Type                | Range | Description                                                        |
| --------------------- | ------------------- | ----- | ------------------------------------------------------------------ |
| `num_images`          | `integer`           | 1–10  | Number of images to generate.                                      |
| `seed`                | `integer`           | —     | Random seed for reproducible results.                              |
| `num_inference_steps` | `integer`           | 1–150 | Denoising steps (FAL models). More steps = higher quality, slower. |
| `guidance_scale`      | `number`            | 0–30  | How closely to follow the prompt (FAL models).                     |
| `temperature`         | `number`            | 0–2   | Sampling temperature for creativity (Vertex models).               |
| `negative_prompt`     | `string` (max 2000) | —     | Text describing elements to exclude from the image.                |

***

### Response Schemas

#### 200 OK — Synchronous Result

Returned when `celery: false` (default). Generation is complete and images are included.

```json
{
  "task_id": "task_1705312200000_abc123",
  "prompt": "A cat on the Eiffel Tower",
  "seed": 42,
  "images": [
    {
      "url": "https://storage.example.com/image.png",
      "width": 1024,
      "height": 1024,
      "content_type": "image/png"
    }
  ],
  "usage": {
    "images_generated": 1,
    "cost_usd": 0.025,
    "cost_credits": 7
  }
}
```

| Field                    | Type            | Description                                                       |
| ------------------------ | --------------- | ----------------------------------------------------------------- |
| `task_id`                | `string`        | Unique identifier for this generation task.                       |
| `prompt`                 | `string`        | The prompt used for generation (echoed back).                     |
| `seed`                   | `integer`       | Seed used (if applicable). Pass back to reproduce the same image. |
| `images`                 | `ImageOutput[]` | Array of generated images.                                        |
| `images[].url`           | `string`        | Public URL of the generated image.                                |
| `images[].width`         | `integer`       | Image width in pixels.                                            |
| `images[].height`        | `integer`       | Image height in pixels.                                           |
| `images[].content_type`  | `string`        | MIME type, e.g. `image/png`.                                      |
| `usage`                  | `object`        | Billing information.                                              |
| `usage.images_generated` | `integer`       | Number of images generated.                                       |
| `usage.cost_usd`         | `number`        | Cost in USD.                                                      |
| `usage.cost_credits`     | `number`        | Cost in platform credits (1 credit = $0.005 USD).                 |

#### 202 Accepted — Async Task Created

Returned when `celery: true`. Poll the task endpoint to retrieve the result.

```json
{
  "task_id": "task_1705312200000_abc123",
  "status": "PENDING",
  "type": "image"
}
```

| Field     | Type     | Description                                     |
| --------- | -------- | ----------------------------------------------- |
| `task_id` | `string` | Use this to poll `GET /api/v1/tasks/{task_id}`. |
| `status`  | `string` | Initial status: always `PENDING`.               |
| `type`    | `string` | Always `image` for image generation tasks.      |

#### Error Responses

| Status | `error` field           | Description                                                |
| ------ | ----------------------- | ---------------------------------------------------------- |
| `400`  | `Bad Request`           | Invalid parameters or unsupported model/field combination. |
| `429`  | `Rate Limited`          | Too many requests. See `retryAfter` field.                 |
| `500`  | `Internal Server Error` | Unexpected server-side failure.                            |

```json
{
  "error": "Bad Request",
  "message": "aspect_ratio '21:9' is not supported by model 'vertex/nano-banana-pro'",
  "retryAfter": null
}
```

***

### Code Examples

#### JavaScript / TypeScript

```typescript
// Synchronous (default)
const response = await fetch('https://api.platform.qolaba.ai/api/v1/images/generate', {
  method: 'POST',
  headers: {
    'X-API-Key': process.env.API_KEY!,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    model: 'vertex/nano-banana-flash',
    prompt: 'A photorealistic cat sitting on the Eiffel Tower at sunset',
    aspect_ratio: '16:9',
    num_images: 1
  }),
});

const result = await response.json();
console.log(result.images[0].url);
```

```typescript
// Asynchronous
const initRes = await fetch('https://api.platform.qolaba.ai/api/v1/images/generate', {
  method: 'POST',
  headers: { 'X-API-Key': process.env.API_KEY!, 'Content-Type': 'application/json' },
  body: JSON.stringify({
    model: 'vertex/nano-banana-flash',
    prompt: 'A highly detailed oil painting of a dragon',
    celery: true,
  }),
});

const { task_id } = await initRes.json(); // 202 Accepted

// Poll until done
const result = await pollTask(task_id);
```

#### cURL

```bash
# Sync - wait for result
curl -X POST https://api.platform.qolaba.ai/api/v1/images/generate \
  -H "X-API-Key: $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "vertex/nano-banana-flash",
    "prompt": "A product photo of wireless headphones",
    "quality": "high",
    "background": "transparent"
  }'

# Async - fire and forget
curl -X POST https://api.platform.qolaba.ai/api/v1/images/generate \
  -H "X-API-Key: $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "vertex/nano-banana-flash",
    "prompt": "Ultra-detailed fantasy landscape",
    "celery": true
  }'
```

***

### Async Polling Flow

When using `celery: true`, poll the task endpoint until `status` is `SUCCESS` or `FAILED`.

#### Poll Endpoint

```
GET /api/v1/tasks/{task_id}
```

#### Task Status Values

| Status       | Description                                               |
| ------------ | --------------------------------------------------------- |
| `PENDING`    | Task is queued, not yet started.                          |
| `PROCESSING` | Generation is in progress.                                |
| `SUCCESS`    | Complete. `output` field contains the images.             |
| `FAILED`     | Generation failed. See `error` and `error_detail` fields. |

#### Polling Example

```typescript
async function pollTask(taskId: string, intervalMs = 2000, timeoutMs = 300000) {
  const deadline = Date.now() + timeoutMs;

  while (Date.now() < deadline) {
    const res = await fetch(`https://api.platform.qolaba.ai/api/v1/tasks/${taskId}`, {
      headers: { 'X-API-Key': process.env.API_KEY! },
    });
    const task = await res.json();

    if (task.status === 'SUCCESS') {
      return task.output; // ImageGenerationResponse
    }

    if (task.status === 'FAILED') {
      throw new Error(`Task failed: ${task.error_detail?.message ?? task.error}`);
    }

    await new Promise(resolve => setTimeout(resolve, intervalMs));
  }

  throw new Error('Task polling timed out');
}
```

#### Task Response Shape

```json
{
  "task_id": "task_1705312200000_abc123",
  "status": "SUCCESS",
  "type": "image",
  "output": {
    "images": [{ "url": "...", "width": 1024, "height": 1024 }],
    "prompt": "...",
    "usage": { "images_generated": 1, "cost_usd": 0.06, "cost_credits": 12 }
  },
  "time_required_ms": 8543,
  "logs": [
    { "timestamp": "2024-01-15T10:30:00Z", "message": "Processing started" },
    { "timestamp": "2024-01-15T10:30:08Z", "message": "Generation complete" }
  ]
}
```

***

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/text-to-image-copy>" %}


# Text-to-Image

Edit, transform, and compose images using Nano Banana Flash and Nano Banana Pro — Google Gemini-powered models that understand both text and images simultaneously.

## Text to image API

`POST /api/v1/images/generate`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

***

### Models at a Glance

| Model ID                   | Name            | Engine           | Speed    | Max Resolution | Max Images | Key Feature                           |
| -------------------------- | --------------- | ---------------- | -------- | -------------- | ---------- | ------------------------------------- |
| `vertex/nano-banana-flash` | Nano Banana 2   | Gemini Flash     | Fast     | 4K             | 10         | Search grounding                      |
| `vertex/nano-banana-pro`   | Nano Banana Pro | Gemini Pro       | Standard | 4K             | 10         | Text rendering, character consistency |
| `vertex/imagen-4`          | Imagen 4        | Imagen 4.0       | Standard | 2K             | 4          | Balanced quality + adherence          |
| `vertex/imagen-4-fast`     | Imagen 4 Fast   | Imagen 4.0 Fast  | Fastest  | \~1K           | 4          | Lowest cost, highest throughput       |
| `vertex/imagen-4-ultra`    | Imagen 4 Ultra  | Imagen 4.0 Ultra | Slowest  | 2K             | 4          | Highest quality, complex prompts      |

***

### Nano Banana Flash — `vertex/nano-banana-flash`

Powered by **Gemini 3.1 Flash**. Best for fast generation, real-time search-grounded accuracy, high-volume workflows, and image-to-image editing.

#### Minimal

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "A cozy Japanese ramen shop on a rainy night, steam rising from bowls, warm ambient lighting"
}
```

#### Standard

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "A hyperrealistic macro photograph of a butterfly on a lavender flower, water droplets on wings, golden hour light",
  "aspect_ratio": "4:3",
  "quality": "2K",
  "num_images": 1,
  "temperature": 1.0,
  "seed": 42
}
```

#### With Search Grounding

Enable real-time Google Search so the model generates factually and visually accurate content — real architecture, real locations, real brand identities.

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "The interior of the Pantheon in Rome with sunlight streaming through the oculus, accurate architecture and proportions",
  "aspect_ratio": "4:3",
  "quality": "2K",
  "use_search_grounding": true
}
```

#### Multiple Variations

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "A minimalist flat illustration of a coffee cup with steam, pastel color palette, icon style",
  "aspect_ratio": "1:1",
  "quality": "1K",
  "num_images": 4,
  "temperature": 1.1
}
```

#### Parameters

| Field                  | Type      | Default      | Options                  | Description                                          |
| ---------------------- | --------- | ------------ | ------------------------ | ---------------------------------------------------- |
| `prompt`               | `string`  | **required** | max 4000 chars           | Text description of the image                        |
| `aspect_ratio`         | `string`  | `1:1`        | See Aspect Ratio Support | Output dimensions                                    |
| `quality`              | `string`  | `2K`         | `512`, `1K`, `2K`, `4K`  | Output resolution. `4K` triggers higher billing rate |
| `num_images`           | `integer` | `1`          | 1–10                     | Variations per call                                  |
| `temperature`          | `number`  | `1.0`        | 0–2                      | Creativity. Lower = faithful, higher = creative      |
| `seed`                 | `integer` | —            | any integer              | Reproducibility. Same seed + prompt = same image     |
| `use_search_grounding` | `boolean` | `false`      | —                        | Real-time Google Search context for accuracy         |
| `celery`               | `boolean` | `false`      | —                        | `true` = async, returns `task_id` immediately        |

***

### Nano Banana Pro — `vertex/nano-banana-pro`

Powered by **Gemini 3 Pro**. Best for production output, multi-language text rendering in images, 4K exports, and consistent character/product depiction across generations.

#### Minimal

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "A luxury watch on a black marble surface, dramatic spotlight lighting, product photography"
}
```

#### Standard

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "A cinema-quality portrait of a woman in a traditional Japanese kimono standing in a cherry blossom garden, shallow depth of field, golden hour lighting",
  "aspect_ratio": "3:4",
  "quality": "4K",
  "num_images": 1,
  "temperature": 0.9,
  "seed": 12345
}
```

#### With Text in Image

Nano Banana Pro renders text in images with high accuracy across multiple languages.

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "A clean coffee shop menu board with the text 'Espresso $3.50', 'Latte $4.50', 'Cappuccino $4.00' written in elegant chalk lettering on a dark background",
  "aspect_ratio": "3:4",
  "quality": "2K",
  "temperature": 0.8
}
```

#### With Reference Images (Character / Product Consistency)

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Generate this character standing in a futuristic neon-lit city plaza at night. Keep exact face, hairstyle, and outfit from the references.",
  "reference_images": [
    {
      "url": "https://your-cdn.com/character-front.jpg",
      "description": "Character front view — match face, hair color, and facial features exactly"
    },
    {
      "url": "https://your-cdn.com/character-outfit.jpg",
      "description": "Character outfit — match the denim jacket, white t-shirt, and black jeans"
    }
  ],
  "quality": "2K",
  "aspect_ratio": "9:16",
  "temperature": 0.85
}
```

#### Parameters

| Field              | Type       | Default      | Options                  | Description                                                                                  |
| ------------------ | ---------- | ------------ | ------------------------ | -------------------------------------------------------------------------------------------- |
| `prompt`           | `string`   | **required** | max 4000 chars           | Text description of the image                                                                |
| `aspect_ratio`     | `string`   | `1:1`        | See Aspect Ratio Support | Output dimensions                                                                            |
| `quality`          | `string`   | `2K`         | `1K`, `2K`, `4K`         | Output resolution. `4K` triggers higher billing rate                                         |
| `num_images`       | `integer`  | `1`          | 1–10                     | Variations per call                                                                          |
| `temperature`      | `number`   | `1.0`        | 0–2                      | Creativity. Lower = faithful, higher = creative                                              |
| `seed`             | `integer`  | —            | any integer              | Reproducibility                                                                              |
| `reference_images` | `object[]` | —            | max 14                   | Visual references for consistency. Each has `url` and optional `description` (max 500 chars) |
| `celery`           | `boolean`  | `false`      | —                        | `true` = async processing                                                                    |

***

### Imagen 4 — `vertex/imagen-4`

Powered by **Imagen 4.0**. Best for general-purpose generation with strict prompt adherence, multilingual support, and configurable safety controls.

#### Minimal

```json
{
  "model": "vertex/imagen-4",
  "prompt": "A hyperrealistic close-up of a dew-covered spider web at dawn, macro photography"
}
```

#### Standard

```json
{
  "model": "vertex/imagen-4",
  "prompt": "A sweeping landscape of the Scottish Highlands at golden hour, rolling heather-covered hills, dramatic storm clouds breaking, a lone castle in the distance",
  "aspect_ratio": "16:9",
  "num_images": 1,
  "image_size": "2K"
}
```

#### With All Safety Controls

```json
{
  "model": "vertex/imagen-4",
  "prompt": "A professional portrait of a businessperson in a modern office, natural window light, confident expression",
  "aspect_ratio": "3:4",
  "image_size": "2K",
  "num_images": 2,
  "output_compression_quality": 90,
  "person_generation": "allow_adult",
  "safety_filter_level": "block_medium_and_above",
  "enhance_prompt": false,
  "add_watermark": false
}
```

***

### Imagen 4 Fast — `vertex/imagen-4-fast`

Powered by **Imagen 4.0 Fast**. Best for bulk generation, rapid prototyping, and cost-sensitive workloads. Fixed to approximately 1K resolution.

#### Minimal

```json
{
  "model": "vertex/imagen-4-fast",
  "prompt": "A vibrant flat illustration of a bicycle in a sunny park, digital art style"
}
```

#### Standard

```json
{
  "model": "vertex/imagen-4-fast",
  "prompt": "A simple icon of a shield with a checkmark, flat design, blue and white, clean vector style",
  "aspect_ratio": "1:1",
  "num_images": 4
}
```

#### Batch Generation

```json
{
  "model": "vertex/imagen-4-fast",
  "prompt": "A colorful abstract geometric background pattern, modern gradient, suitable for mobile app splash screen",
  "aspect_ratio": "9:16",
  "num_images": 4,
  "celery": true
}
```

***

### Imagen 4 Ultra — `vertex/imagen-4-ultra`

Powered by **Imagen 4.0 Ultra**. Best for the highest quality outputs, complex multi-element compositions, and strict instruction following. Use when quality is more important than speed or cost.

#### Minimal

```json
{
  "model": "vertex/imagen-4-ultra",
  "prompt": "A breathtakingly detailed oil painting of a medieval port city at twilight, torchlit cobblestone streets, galleons in harbor, figures in period clothing"
}
```

#### Standard

```json
{
  "model": "vertex/imagen-4-ultra",
  "prompt": "A hyperrealistic underwater scene: a coral reef ecosystem at depth, shafts of turquoise sunlight filtering through the surface, diverse marine life including clownfish, sea turtles, and manta rays, razor-sharp focus throughout the frame",
  "aspect_ratio": "16:9",
  "image_size": "2K",
  "num_images": 1
}
```

#### Maximum Quality

```json
{
  "model": "vertex/imagen-4-ultra",
  "prompt": "A sweeping aerial panorama of the Dolomites in Italy at sunrise, snow-capped peaks glowing pink and orange, deep valleys in morning shadow, a winding mountain road visible far below, cinematic 8K quality",
  "aspect_ratio": "16:9",
  "image_size": "2K",
  "num_images": 1,
  "enhance_prompt": false,
  "add_watermark": false,
  "safety_filter_level": "block_medium_and_above"
}
```

***

### Aspect Ratio Support

| Ratio  | NB Flash | NB Pro | Imagen 4 / Fast / Ultra |
| ------ | :------: | :----: | :---------------------: |
| `1:1`  |    Yes   |   Yes  |           Yes           |
| `4:3`  |    Yes   |   Yes  |           Yes           |
| `3:4`  |    Yes   |   Yes  |           Yes           |
| `16:9` |    Yes   |   Yes  |           Yes           |
| `9:16` |    Yes   |   Yes  |           Yes           |
| `3:2`  |    Yes   |   Yes  |            —            |
| `2:3`  |    Yes   |   Yes  |            —            |
| `21:9` |    Yes   |   Yes  |            —            |
| `1:4`  |    Yes   |   Yes  |            —            |
| `4:1`  |    Yes   |   Yes  |            —            |
| `5:4`  |    Yes   |   Yes  |            —            |
| `4:5`  |    Yes   |   Yes  |            —            |


# Image-to-Image

Edit, transform, and compose images using Nano Banana Flash and Nano Banana Pro — Google Gemini-powered models that understand both text and images simultaneously.

Image-to-Image  API

`POST /api/v1/images/generate`

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

### Model Overview

|                                      | `vertex/nano-banana-flash`     | `vertex/nano-banana-pro`   |
| ------------------------------------ | ------------------------------ | -------------------------- |
| **Underlying model**                 | gemini-3.1-flash-image-preview | gemini-3-pro-image-preview |
| **Provider**                         | Google Gemini / Vertex AI      | Google Gemini / Vertex AI  |
| **Speed**                            | Faster                         | Slower                     |
| **Quality**                          | High                           | Excellent                  |
| **Max resolution**                   | 4K                             | 4K                         |
| **Quality tiers**                    | `512`, `1K`, `2K`, `4K`        | `1K`, `2K`, `4K`           |
| **Max source images** (`image_urls`) | 14                             | 14                         |
| **Max reference images**             | 14                             | 14                         |
| **Search grounding**                 | Yes                            | No                         |
| **Text rendering**                   | Good                           | Excellent                  |
| **Subject consistency**              | Good                           | Excellent                  |
| **Cost per image (2K)**              | \~$0.08                        | \~$0.161                   |

***

### Capabilities

Both models support the following image-to-image operations:

| Capability                      | Description                                                       |
| ------------------------------- | ----------------------------------------------------------------- |
| **Image editing**               | Modify specific parts of an image based on a text instruction     |
| **Style transfer**              | Apply the visual style of one image to another                    |
| **Background replacement**      | Swap the background while keeping the subject                     |
| **Multi-image composition**     | Combine elements from multiple source images into one             |
| **Reference-based consistency** | Keep a character, product, or style consistent across generations |
| **Variations**                  | Generate alternative versions of an existing image                |
| **Object placement**            | Insert a logo, product, or object into a scene                    |
| **Appearance change**           | Change clothing, color, texture, or features                      |
| **Search-grounded editing**     | Edit using real-world knowledge (Flash only)                      |

***

### Request Reference

#### Required Fields

| Field        | Type                | Description                                            |
| ------------ | ------------------- | ------------------------------------------------------ |
| `model`      | `string`            | `vertex/nano-banana-flash` or `vertex/nano-banana-pro` |
| `prompt`     | `string` (max 4000) | Editing instruction or description.                    |
| `image_urls` | `string[]`          | One or more publicly accessible source image URLs.     |

#### Valid `aspect_ratio` Values

```
1:1   1:4   1:8   2:3   3:2   3:4   4:1   4:3   4:5   5:4   8:1   9:16   16:9   21:9
```

***

### Use Cases & Payloads

***

#### 1. Edit a Specific Part of an Image

Change a targeted element while leaving everything else untouched.

**Flash — fast iteration:**

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Change the car color to matte black, keep everything else exactly the same",
  "image_urls": ["https://cdn.qolaba.ai/1778648845113_9i9sl6e255.jpeg"],
  "quality": "2K",
  "aspect_ratio": "16:9",
  "temperature": 0.8
}
```

**Pro — production output:**

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Change the color of the sofa to deep emerald green, keep the rest of the room unchanged",
  "image_urls": ["https://cdn.qolaba.ai/1778648845113_9i9sl6e255.jpeg"],
  "quality": "4K",
  "aspect_ratio": "16:9",
  "temperature": 0.75
}
```

***

#### 2. Style Transfer

Apply the visual style, mood, or artistic look of a reference image to a source image.

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Redraw this photo in the style of the reference image — apply the same painterly brush strokes, warm color palette, and impressionist lighting. Preserve the original composition and subject.",
  "image_urls": ["https://cdn.qolaba.ai/1778648845113_9i9sl6e255.jpeg"],
  "reference_images": [
    {
      "url": "https://your-cdn.com/van-gogh-style.jpg",
      "description": "Artistic style to apply — impressionist oil painting with swirling brushwork"
    }
  ],
  "quality": "2K",
  "aspect_ratio": "1:1",
  "temperature": 1.1
}
```

**Text-only style transfer (no reference image):**

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Convert this photo into a watercolor painting with soft edges, muted pastel tones, and visible paper texture. Keep the original subject and composition.",
  "image_urls": ["https://cdn.qolaba.ai/1778648845113_9i9sl6e255.jpeg"],
  "quality": "2K",
  "aspect_ratio": "16:9",
  "temperature": 1.0
}
```

***

#### 3. Replace Background

See the dedicated Replace Background guide for full documentation.

**Quick reference:**

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Replace the background with a modern city skyline at golden hour, keep the subject exactly as-is",
  "image_urls": ["https://cdn.qolaba.ai/1778648845113_9i9sl6e255.jpeg"],
  "quality": "2K",
  "aspect_ratio": "16:9"
}
```

***

#### 4. Multi-Image Composition

Combine elements from multiple source images into a single coherent output.

**Merge two scenes:**

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Combine these two images: place the person from the first image into the environment shown in the second image. Match lighting and perspective.",
  "image_urls": [
    "https://cdn.qolaba.ai/1778648845113_9i9sl6e255.jpeg",
    "https://cdn.qolaba.app/1778662155240_ydyuksjbp2b.jpeg"
  ],
  "quality": "2K",
  "aspect_ratio": "16:9",
  "temperature": 0.9
}
```

**Extract and composite a product:**

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Take the bottle from the first image and place it on the table in the second image. Match the lighting, shadow, and perspective of the scene.",
   "image_urls": [
    "https://cdn.qolaba.ai/1778648845113_9i9sl6e255.jpeg",
    "https://cdn.qolaba.app/1778662155240_ydyuksjbp2b.jpeg"
  ],
  "quality": "4K",
  "aspect_ratio": "1:1",
  "temperature": 0.85
}
```

***

#### 5. Character / Product Consistency with References

Keep a specific subject (person, character, product) visually consistent across multiple generations using `reference_images`.

**Consistent character across scenes:**

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Generate an image of this character sitting in a cozy coffee shop reading a book. Maintain exact visual consistency — same face, hairstyle, clothing, and proportions as shown in the reference images.",
  "image_urls": [],
  "reference_images": [
    {
      "url": "https://your-cdn.com/character-front.jpg",
      "description": "Character reference — front view"
    },
    {
      "url": "https://your-cdn.com/character-side.jpg",
      "description": "Character reference — side view"
    },
    {
      "url": "https://your-cdn.com/character-detail.jpg",
      "description": "Character reference — clothing and accessory detail"
    }
  ],
  "quality": "2K",
  "aspect_ratio": "4:3",
  "temperature": 0.85
}
```

**Consistent product in a lifestyle scene:**

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Show this exact product placed on a rustic wooden kitchen counter with morning sunlight. Keep every product detail — shape, label, color, and texture — identical to the reference.",
  "image_urls": [],
  "reference_images": [
    {
      "url": "https://your-cdn.com/product-front.jpg",
      "description": "Product front view reference"
    },
    {
      "url": "https://your-cdn.com/product-label.jpg",
      "description": "Product label and branding detail reference"
    }
  ],
  "quality": "4K",
  "aspect_ratio": "16:9",
  "temperature": 0.8
}
```

***

#### 6. Image Upscale & Re-render

Re-render a low-quality or low-resolution image at higher quality with enhanced detail.

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Re-render this image at high resolution. Enhance sharpness, add fine detail, improve lighting quality, and increase overall visual fidelity. Do not change the composition, colors, or content.",
  "image_urls": ["https://your-cdn.com/low-res.jpg"],
  "quality": "4K",
  "aspect_ratio": "1:1",
  "temperature": 0.7
}
```

**Upscale with artistic enhancement:**

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Upscale this image and enhance the details. Make textures more realistic, sharpen edges, and improve the lighting while keeping the exact same scene and composition.",
  "image_urls": ["https://your-cdn.com/draft.jpg"],
  "quality": "4K",
  "aspect_ratio": "16:9",
  "temperature": 0.75
}
```

***

#### 7. Outfit & Appearance Change

Modify what a subject is wearing or change their visual appearance.

**Change outfit:**

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Change the person's outfit to a formal navy blue business suit with a white shirt and tie. Keep their face, hairstyle, body, and the background exactly the same.",
  "image_urls": ["https://your-cdn.com/person.jpg"],
  "quality": "2K",
  "aspect_ratio": "3:4",
  "temperature": 0.8
}
```

**Change hair color:**

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Change the person's hair color to platinum blonde. Keep everything else — face, clothing, background, and pose — identical.",
  "image_urls": ["https://your-cdn.com/portrait.jpg"],
  "quality": "2K",
  "aspect_ratio": "1:1",
  "temperature": 0.75
}
```

**Product color variant:**

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Change the color of the sneakers to all-black with a white sole. Keep the exact same shoe design, angle, lighting, and background.",
  "image_urls": ["https://your-cdn.com/sneaker-white.jpg"],
  "quality": "2K",
  "aspect_ratio": "1:1",
  "temperature": 0.7,
  "num_images": 4
}
```

***

#### 8. Logo / Object Placement into a Scene

Insert a product, logo, or object into an existing scene naturally.

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Place the logo from the reference image on the front of the white t-shirt worn by the person. The logo should be centered on the chest, sized proportionally, and match the fabric texture and lighting of the shirt.",
  "image_urls": ["https://your-cdn.com/person-white-tshirt.jpg"],
  "reference_images": [
    {
      "url": "https://your-cdn.com/brand-logo.png",
      "description": "Brand logo to place on the t-shirt"
    }
  ],
  "quality": "2K",
  "aspect_ratio": "3:4",
  "temperature": 0.8
}
```

**Product mockup on billboard:**

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Place this product image on the blank billboard in the urban street photo. The product image should fill the billboard naturally, matching the perspective and lighting of the scene.",
  "image_urls": [
    "https://your-cdn.com/street-with-billboard.jpg",
    "https://your-cdn.com/product-ad.jpg"
  ],
  "quality": "2K",
  "aspect_ratio": "16:9",
  "temperature": 0.8
}
```

***

### Prompt Writing Guide

#### State the edit clearly upfront

Start the prompt with what you want to change before describing the result.

```
# Clear
"Change the jacket color to red. Keep the person's face, body, and the background unchanged."

# Unclear
"A person in a red jacket in a park on a sunny day"
```

#### Explicitly protect what should not change

The model won't know what to preserve unless you say so.

```
# Safer
"Change the table material to marble. Keep the objects on the table, the room, the lighting, and the camera angle exactly the same."

# Risky
"A marble table in a living room"
```

#### Separate the subject from the edit

Describe the original subject, then describe the change.

```
"This is a product photo of a glass perfume bottle on a white surface.
Change the background surface from white to dark polished obsidian.
Keep the bottle, its reflection, and the overhead lighting unchanged."
```


# Models

Retrieve all available image generation models.

## Images — Models API

`GET /api/v1/images/models`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

### Request

No request body or query parameters required.

```http
GET /api/v1/images/models
Authorization: Bearer <YOUR_API_KEY>
```

***

### Response

**`200 OK`**

Returns an object containing a `models` array.

```json
{
  "models": [ ... ]
}
```

#### Model Object

| Field                 | Type       | Description                                                                                             |
| --------------------- | ---------- | ------------------------------------------------------------------------------------------------------- |
| `id`                  | `string`   | Unique model identifier. Use this as the `model` field when calling the generate endpoint.              |
| `name`                | `string`   | Human-readable model name.                                                                              |
| `type`                | `string`   | Always `"image"` for image models.                                                                      |
| `provider`            | `string`   | Internal provider key.                                                                                  |
| `providerName`        | `string`   | Display name of the underlying provider (e.g. `"Google Vertex AI"`).                                    |
| `image`               | `string`   | URL of the provider logo.                                                                               |
| `description`         | `string`   | Summary of the model's strengths and ideal use cases.                                                   |
| `detailedDescription` | `string`   | One-line technical summary.                                                                             |
| `capabilities`        | `string[]` | List of supported capabilities. See Capabilities.                                                       |
| `imageSize`           | `string[]` | Supported aspect ratios (e.g. `"16:9"`, `"1:1"`).                                                       |
| `quality`             | `string[]` | Supported output quality tiers (e.g. `"512"`, `"1K"`, `"2K"`, `"4K"`). Not present on all models.       |
| `quality_pricing`     | `object`   | Cost per image in USD for each quality tier. Keys match the `quality` array. Not present on all models. |
| `outputFormat`        | `string[]` | Supported output MIME types (e.g. `"image/png"`, `"image/jpeg"`). Not present on all models.            |
| `maxNumImages`        | `integer`  | Maximum number of images per request.                                                                   |
| `defaultNumImages`    | `integer`  | Default number of images when `num_images` is not specified.                                            |
| `minTemperature`      | `number`   | Minimum allowed `temperature` value. Not present on all models.                                         |
| `maxTemperature`      | `number`   | Maximum allowed `temperature` value. Not present on all models.                                         |
| `defaultTemperature`  | `number`   | Default `temperature` value. Not present on all models.                                                 |
| `maxReferenceImages`  | `integer`  | Maximum number of reference images per request. Not present on all models.                              |
| `supportsSeed`        | `boolean`  | Whether the model accepts a `seed` for reproducible outputs. Assume `false` if absent.                  |

***

### Capabilities

| Value              | Description                                                             |
| ------------------ | ----------------------------------------------------------------------- |
| `text_to_image`    | Generate an image from a text prompt.                                   |
| `image_to_image`   | Transform or edit an existing image guided by a prompt.                 |
| `reference_images` | Use one or more reference images to guide style or subject consistency. |
| `search_grounding` | Augment generation with real-world search context.                      |

***

### Available Models

| Model ID                   | Name            | Provider         | Capabilities                                                            |
| -------------------------- | --------------- | ---------------- | ----------------------------------------------------------------------- |
| `vertex/nano-banana-flash` | Nano Banana 2   | Google Gemini    | text\_to\_image, image\_to\_image, reference\_images, search\_grounding |
| `vertex/nano-banana-pro`   | Nano Banana Pro | Google Gemini    | text\_to\_image, image\_to\_image, reference\_images                    |
| `vertex/imagen-4`          | Imagen 4        | Google Vertex AI | text\_to\_image                                                         |
| `vertex/imagen-4-fast`     | Imagen 4 Fast   | Google Vertex AI | text\_to\_image                                                         |
| `vertex/imagen-4-ultra`    | Imagen 4 Ultra  | Google Vertex AI | text\_to\_image                                                         |

***

### Error Responses

All errors follow this shape:

```json
{
  "error": {
    "message": "string",
    "type": "string",
    "param": null,
    "code": "string"
  }
}
```

| Status | `code`           | Description                                         |
| ------ | ---------------- | --------------------------------------------------- |
| `500`  | `internal_error` | Unexpected server error.                            |
| `502`  | `upstream_error` | Failed to reach the upstream provider.              |
| `504`  | `timeout`        | Request to upstream provider timed out (30s limit). |

***

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/images-models>" %}


# ControlNet

Our API currently supports a variety of 3 ControlNet models, categorized into distinct groups. These include Canny, Depth, and Pose Estimation models, all of which are based on the SDXL framework.

For more detailed information about these ControlNet models, please refer to the [Broken mention](broken://pages/PTC8Od5374PscvfwmZJK) section.

## ControlNet API

<mark style="color:green;">`POST`</mark> `/getImagetoImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name                  | Type   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                |
| --------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| app\_id               | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                                                                                                                                               |
| image                 | string | <p>-> The <code>image</code> parameter specifies the URL of an existing image that will be used as a reference for the generation process. </p><p>-> If the original image dimensions exceed 1536x1536 pixels, the image will be adjusted to fit within this size while preserving the original aspect ratio.</p>                                                                                                                          |
| prompt                | string | -> The `prompt` parameter is the textual input that guides the image generation process. This prompt serves as an artistic compass, shaping the visual output.                                                                                                                                                                                                                                                                             |
| guidance\_scale       | float  | <p>-> The <code>guidance\_scale</code> parameter determines how closely the generated image adheres to the provided prompt. Higher values result in the model following the prompt more closely, while lower values allow for more creative deviation.</p><p>-> The valid range for the <code>guidance\_scale</code> parameter is between 1 and 30.</p>                                                                                    |
| batch                 | int    | <p>-> The <code>batch</code> parameter allows you to specify the number of images to generate at once. </p><p>-> The valid range for this parameter is between 1 and 8.</p>                                                                                                                                                                                                                                                                |
| strength              | float  | <p>-> The <code>strength</code> parameter specifies the degree of transformation applied to the reference image. </p><p>-> A higher <code>strength</code> value (up to 1) results in the generated image closely following the initial reference image. </p>                                                                                                                                                                               |
| negative\_prompt      | string | -> The `negative_prompt` parameter allows you to specify content that you want the image generation model to avoid or minimize in the output. This can be useful for excluding certain visual elements or styles that you do not want to be present in the generated image.                                                                                                                                                                |
| num\_inference\_steps | int    | <p>-> The <code>num\_inference\_steps</code> parameter represents the number of denoising iterations to perform during the image generation process. Generally, more iterations can result in higher-quality images, but they also increase the time required for generation.</p><p>-> The valid range for the <code>num\_inference\_steps</code> parameter is between 1 and 50.</p>                                                       |
| celery                | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.                                                                                                                                    |
| inference\_type       | string | <p>-> The <code>inference\_type</code> parameter allows you to specify the GPU to be used for the image generation task. The supported values are:</p><ul><li><code>a10g</code></li><li><code>a100</code></li><li><code>h100</code></li></ul><p>The different GPU options provide varying levels of performance and capabilities, allowing you to choose the most suitable GPU based on your requirements and the demand for the task.</p> |

**APP IDs for different ControlNet**

<table><thead><tr><th width="392">App ID</th><th>Model Name</th></tr></thead><tbody><tr><td>ap-1us0FK21Ach6eiWxo22is8</td><td>Canny ControlNet</td></tr><tr><td>ap-WrXnJBXy23XpPh6IlH5tRX</td><td>Depth ControlNet</td></tr><tr><td>ap-a1b2c3d4e5f6g7h8i9j0kq</td><td>Pose ControlNet</td></tr></tbody></table>

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/controlnet>" %}


# Inpainting

For more detailed information on the workings of this model, please refer to the [Broken mention](broken://pages/sjqlz3z6PT6o9fLzyVR5) page.

## Inpainting API

<mark style="color:green;">`POST`</mark> `/getImagetoImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name             | Type   | Description                                                                                                                                                                                                                                                                                                                                                                          |
| ---------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| app\_id          | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                                                                                         |
| image            | string | <p>-> The <code>image</code> parameter specifies the URL of an existing image that will be used as a reference for the generation process. </p><p>-> If the original image dimensions exceed 1536x1536 pixels, the image will be adjusted to fit within this size while preserving the original aspect ratio.</p>                                                                    |
| mask\_url        | string | <p>-> The <code>mask\_url</code> parameter is used for providing image mask. It specifies a binary mask image that defines the regions of the input image that should be inpainted.</p><p>-> The mask image should have the same dimensions as the input image, with white pixels indicating the areas to be inpainted, and black pixels representing the areas to be preserved.</p> |
| prompt           | string | -> The `prompt` parameter is the textual input that guides the image generation process. This prompt serves as an artistic compass, shaping the visual output.                                                                                                                                                                                                                       |
| batch            | int    | <p>-> The <code>batch</code> parameter allows you to specify the number of images to generate at once. </p><p>-> The valid range for this parameter is between 1 and 8.</p>                                                                                                                                                                                                          |
| negative\_prompt | string | -> The `negative_prompt` parameter allows you to specify content that you want the image generation model to avoid or minimize in the output. This can be useful for excluding certain visual elements or styles that you do not want to be present in the generated image.                                                                                                          |
| celery           | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.                                                                              |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/inpainting>" %}


# Replace Background

Replace or swap the background of any image while preserving the foreground subject using the Image Generation API.

***

### Overview

The Replace Background feature takes an existing photo and rewrites the background based on your text prompt, while keeping the foreground subject intact.

## Replace Background API

`POST /api/v1/images/generate`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**How it works:**

```
  Input photo + prompt
        │
        ▼
  [Gemini model understands
   subject vs. background]
        │
        ▼
  New background rendered
  around the subject
        │
        ▼
  Output image returned
```

**Two approaches:**

| Approach                      | When to Use                                   | Models                                               |
| ----------------------------- | --------------------------------------------- | ---------------------------------------------------- |
| **Direct replacement**        | Background is clearly distinct from subject   | `vertex/nano-banana-flash`, `vertex/nano-banana-pro` |
| **Cutout + replace pipeline** | Hair, fur, complex edges, transparent objects | BiRefNet → Nano Banana                               |

***

### Supported Models

Only Gemini-based models support image-to-image editing. Imagen 4 variants are **text-to-image only** and do not accept source images.

| Model                      | Speed    | Max Resolution | Max Source Images | Best For                                        |
| -------------------------- | -------- | -------------- | ----------------- | ----------------------------------------------- |
| `vertex/nano-banana-flash` | Fast     | 4K             | 14                | Rapid iteration, prototyping, batch jobs        |
| `vertex/nano-banana-pro`   | Standard | 4K             | 14                | Production, text in images, precise consistency |

***

### Quick Start

```bash
curl -X POST https://api.platform.qolaba.ai/api/v1/images/generate \
  -H "X-API-Key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "vertex/nano-banana-flash",
    "prompt": "Replace the background with a modern city skyline at golden hour, keep the subject exactly as-is",
    "image_urls": ["https://cdn.qolaba.ai/1778648245451_d2c9wj0yf2a.jpeg"],
    "quality": "1K",
    "aspect_ratio": "16:9"
  }'
```

**Response `200 OK`:**

```json
{
  "task_id": "task_1705312200000_abc123",
  "images": [
    {
      "url": "https://storage.example.com/result.png",
      "width": 2048,
      "height": 1152,
      "content_type": "image/png"
    }
  ],
  "usage": {
    "images_generated": 1,
    "cost_usd": 0.08,
    "cost_credits": 16
  }
}
```

***

### Request Reference

#### Required Fields

| Field        | Type                | Description                                            |
| ------------ | ------------------- | ------------------------------------------------------ |
| `model`      | `string`            | `vertex/nano-banana-flash` or `vertex/nano-banana-pro` |
| `prompt`     | `string` (max 4000) | Describe the new background. See Prompt Writing Guide. |
| `image_urls` | `string[]`          | Array containing the URL of the source image to edit.  |

#### Optional Fields

| Field                  | Type       | Default | Description                                                                 |
| ---------------------- | ---------- | ------- | --------------------------------------------------------------------------- |
| `quality`              | `string`   | `2K`    | Output resolution. Flash: `512 \| 1K \| 2K \| 4K`. Pro: `1K \| 2K \| 4K`.   |
| `aspect_ratio`         | `string`   | `1:1`   | Output dimensions. See valid values below.                                  |
| `num_images`           | `integer`  | `1`     | Number of variations to generate (1–4).                                     |
| `temperature`          | `number`   | `1.0`   | Creativity level. `0.7–0.9` = faithful, `1.2–1.5` = more creative.          |
| `reference_images`     | `object[]` | —       | Up to 14 reference images for style/environment matching.                   |
| `use_search_grounding` | `boolean`  | `false` | Enable real-time search for factually accurate backgrounds. **Flash only.** |
| `celery`               | `boolean`  | `false` | `true` = async processing, returns `task_id` immediately.                   |
| `call_webhook`         | `boolean`  | `false` | Send a POST callback when generation completes.                             |
| `webhook_url`          | `string`   | —       | Webhook destination URL. Required when `call_webhook: true`.                |
| `webhook_secret`       | `string`   | —       | Secret sent as `x-webhook-secret` header in the callback.                   |

#### Valid `aspect_ratio` Values

```
1:1   3:2   2:3   4:3   3:4   16:9   9:16   4:5   5:4   21:9
```

***

### Payloads by Use Case

#### 1. Simple Background Swap

Replace a busy background with a clean, minimal one.

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Replace the background with a plain white studio backdrop with soft even lighting, keep the subject unchanged",
  "image_urls": ["https://your-cdn.com/product-photo.jpg"],
  "quality": "2K",
  "aspect_ratio": "1:1"
}
```

***

#### 2. Outdoor / Scenic Background

Place the subject in a natural or architectural environment.

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Replace the background with a lush green forest at sunrise with soft fog, preserve the person in the foreground exactly",
  "image_urls": ["https://your-cdn.com/portrait.jpg"],
  "quality": "4K",
  "aspect_ratio": "9:16",
  "temperature": 0.9
}
```

***

#### 3. Product Photography Background

Ideal for e-commerce — clean, professional backgrounds.

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Replace the background with a light grey gradient studio background with subtle shadow underneath the product, keep the product exactly as-is",
  "image_urls": ["https://your-cdn.com/product.jpg"],
  "quality": "2K",
  "aspect_ratio": "1:1",
  "temperature": 0.7,
  "num_images": 3
}
```

***

#### 4. Background Matching a Reference Image

Use a reference photo to match a specific environment or style.

```json
{
  "model": "vertex/nano-banana-pro",
  "prompt": "Swap the background to closely match the environment in the reference image, preserve the foreground subject without any changes",
  "image_urls": ["https://your-cdn.com/subject.jpg"],
  "reference_images": [
    {
      "url": "https://your-cdn.com/target-environment.jpg",
      "description": "Use this as the target background scene and lighting"
    }
  ],
  "quality": "2K",
  "aspect_ratio": "16:9",
  "temperature": 0.85
}
```

***

#### 5. Virtual Meeting / Office Background

Professional background replacement for profile photos or video thumbnails.

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Replace the background with a modern open-plan office with blurred depth-of-field effect, keep the person in the foreground sharp and unchanged",
  "image_urls": ["https://your-cdn.com/headshot.jpg"],
  "quality": "2K",
  "aspect_ratio": "16:9",
  "temperature": 1.0
}
```

***

#### 6. Real-World Context with Search Grounding

Generate a geographically or factually accurate background using live search data. **Nano Banana Flash only.**

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Replace the background with the interior of the Louvre Museum in Paris near the Mona Lisa, keep the subject as-is",
  "image_urls": ["https://your-cdn.com/tourist.jpg"],
  "quality": "2K",
  "aspect_ratio": "4:3",
  "use_search_grounding": true
}
```

***

#### 7. Multiple Variations in One Request

Generate several background options in a single call.

```json
{
  "model": "vertex/nano-banana-flash",
  "prompt": "Replace the background with a vibrant night city street with neon lights and rain reflections, keep the subject exactly as-is",
  "image_urls": ["https://your-cdn.com/portrait.jpg"],
  "quality": "2K",
  "aspect_ratio": "9:16",
  "num_images": 4,
  "temperature": 1.2
}
```

***

### Prompt Writing Guide

The prompt is the most important factor in getting clean background replacement. Follow these principles:

#### Always Anchor the Subject

Include an explicit instruction to preserve the foreground. Without it the model may alter both subject and background.

```
# Good
"Replace the background with a tropical beach at sunset, keep the person in the foreground exactly as-is"

# Risky — may alter the subject too
"A person standing on a tropical beach at sunset"
```

#### Describe the New Background Specifically

Vague prompts lead to inconsistent results. Be specific about lighting, distance, depth of field, and mood.

```
# Vague
"outdoor background"

# Specific
"outdoor park background with soft bokeh blur, overcast natural lighting, green trees in the distance"
```

#### Specify Lighting to Match the Subject

If the source photo has directional light, describe matching light in the background to avoid a compositing-look.

```
"Replace background with a sunset beach, warm golden hour lighting from the right side to match the subject's lighting"
```

#### Use Negative Instructions When Needed

State what you do NOT want to help the model avoid common pitfalls.

```
"Replace the background with a snowy mountain scene. Do not change the subject's clothing, face, or body. No watermarks. No text overlay."
```

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/replace-background>" %}


# Face Consistency

The "Consistent Face Model" allows you to generate a series of images where the face maintains a consistent appearance across the generated outputs.

To use this feature, you provide an initial facial image as a reference. The model then uses this reference to create new images while keeping the facial features aligned with the original picture. This ensures that the face structure and characteristics remain faithful to your initial prompt.

For a more detailed explanation of how the Consistent Face Model works, please refer to the [Broken mention](broken://pages/yScOTdRur41AMVQlZXFt) page.

## Face consistency API

<mark style="color:green;">`POST`</mark> `/getImagetoImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name                  | Type   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                |
| --------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| app\_id               | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                                                                                                                                               |
| image                 | string | <p>-> The <code>image</code> parameter specifies the URL of an existing image that will be used as a reference for the generation process. </p><p>-> If the original image dimensions exceed 1536x1536 pixels, the image will be adjusted to fit within this size while preserving the original aspect ratio.</p>                                                                                                                          |
| prompt                | string | -> The `prompt` parameter is the textual input that guides the image generation process. This prompt serves as an artistic compass, shaping the visual output.                                                                                                                                                                                                                                                                             |
| batch                 | int    | <p>-> The <code>batch</code> parameter allows you to specify the number of images to generate at once. </p><p>-> The valid range for this parameter is between 1 and 8.</p>                                                                                                                                                                                                                                                                |
| height                | int    | <p>-> The <code>height</code> parameter represents the vertical dimension of an image. </p><p>-> The valid range for the parameter is between 256 and 1536 pixels. </p>                                                                                                                                                                                                                                                                    |
| width                 | int    | <p>-> The <code>width</code> parameter represents the horizontal dimension of an image. </p><p>-> The valid range for the parameter is between 256 and 1536 pixels. </p>                                                                                                                                                                                                                                                                   |
| num\_inference\_steps | int    | <p>-> The <code>num\_inference\_steps</code> parameter represents the number of denoising iterations to perform during the image generation process. Generally, more iterations can result in higher-quality images, but they also increase the time required for generation.</p><p>-> The valid range for the <code>num\_inference\_steps</code> parameter is between 1 and 50.</p>                                                       |
| guidance\_scale       | float  | <p>-> The <code>guidance\_scale</code> parameter determines how closely the generated image adheres to the provided prompt. Higher values result in the model following the prompt more closely, while lower values allow for more creative deviation.</p><p>-> The valid range for the <code>guidance\_scale</code> parameter is between 1 and 30.</p>                                                                                    |
| negative\_prompt      | string | -> The `negative_prompt` parameter allows you to specify content that you want the image generation model to avoid or minimize in the output. This can be useful for excluding certain visual elements or styles that you do not want to be present in the generated image.                                                                                                                                                                |
| strength              | float  | <p>-> The <code>strength</code> parameter specifies the degree of transformation applied to the reference image. </p><p>-> A higher <code>strength</code> value (up to 1) results in the generated image closely following the reference image.</p>                                                                                                                                                                                        |
| celery                | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.                                                                                                                                    |
| inference\_type       | string | <p>-> The <code>inference\_type</code> parameter allows you to specify the GPU to be used for the image generation task. The supported values are:</p><ul><li><code>a10g</code></li><li><code>a100</code></li><li><code>h100</code></li></ul><p>The different GPU options provide varying levels of performance and capabilities, allowing you to choose the most suitable GPU based on your requirements and the demand for the task.</p> |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/face-consistency>" %}


# Face Avatar

The concept of a "Face Avatar" involves the creation of images where the face maintains a consistent appearance across multiple generated Images. To achieve this, you start with an initial facial image and provide it as a reference. Using this reference, a sophisticated model goes to work, crafting new images while keeping the facial features in line with the original picture. In essence, the model uses your input to generate images where the face remains faithful to the structure and characteristics you've specified in your prompt. This way, you can effortlessly create a series of images with a consistent and recognizable face.

Working insights :

In this model, the default prompt settings includes a prompt, “SFW Content, black plain background, joyous dating profile, VECTOR CARTOON ILLUSTRATION, half-body shot portrait enjoyable pleasing pleasurable nice {gender\_word}, looking at camera, Relaxed, Charming, Cordial, Gracious, 5 o clock shadow, 3d bitmoji avatar render, pixar, high def textures 8k, highly detailed, 3d render, award winning, no background elements".

If you provide a custom prompt, it will override these default settings. To include a specific gender in your custom prompt, you must explicitly mention it in custom prompt. If you prefer to use the default settings without specifying any custom prompt, simply pass `None` in the prompt parameter. Subsequently, you can define the gender by using the gender parameter, which will apply the specified gender to the default image configuration.&#x20;

## Face Avatar API

<mark style="color:green;">`POST`</mark> `/getImageToImage`

**Headers**

<table><thead><tr><th width="380">Name</th><th>Value</th></tr></thead><tbody><tr><td>Content-Type</td><td><code>application/json</code></td></tr><tr><td>Authorization</td><td><code>Bearer &#x3C;token></code></td></tr></tbody></table>

**Body**

| Name                  | Type   | Description                                                                                                                                                                                                                                                                                                                                                                          |
| --------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| app\_id               | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                                                                                         |
| image                 | string | <p>-> The <code>file\_url</code> parameter specifies the URL of an existing image that will be used as a reference for the generation process. </p><p>-> If the original image dimensions exceed 1536x1536 pixels, the image will be adjusted to fit within this size while preserving the original aspect ratio.</p>                                                                |
| prompt                | string | -> The `prompt` parameter is the textual input that guides the image generation process. This prompt serves as an artistic compass, shaping the visual output.                                                                                                                                                                                                                       |
| batch                 | int    | <p>-> The <code>batch</code> parameter allows you to specify the number of images to generate at once. </p><p>-> The valid range for this parameter is between 1 and 8.</p>                                                                                                                                                                                                          |
| height                | int    | <p>-> The <code>height</code> parameter represents the vertical dimension of an image. </p><p>-> The valid range for the parameter is between 256 and 1536 pixels. </p>                                                                                                                                                                                                              |
| width                 | int    | <p>-> The <code>width</code> parameter represents the horizontal dimension of an image. </p><p>-> The valid range for the parameter is between 256 and 1536 pixels. </p>                                                                                                                                                                                                             |
| celery                | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.                                                                              |
| num\_inference\_steps | int    | <p>-> The <code>num\_inference\_steps</code> parameter represents the number of denoising iterations to perform during the image generation process. Generally, more iterations can result in higher-quality images, but they also increase the time required for generation.</p><p>-> The valid range for the <code>num\_inference\_steps</code> parameter is between 1 and 50.</p> |
| guidance\_scale       | float  | <p>-> The <code>guidance\_scale</code> parameter determines how closely the generated image adheres to the provided prompt. Higher values result in the model following the prompt more closely, while lower values allow for more creative deviation.</p><p>-> The valid range for the <code>guidance\_scale</code> parameter is between 1 and 30.</p>                              |
| negative\_prompt      | string | -> The `negative_prompt` parameter allows you to specify content that you want the image generation model to avoid or minimize in the output. This can be useful for excluding certain visual elements or styles that you do not want to be present in the generated image.                                                                                                          |
| gender                | string | -> The `gender` parameter is used to guide the image generation model towards producing outputs with a specific gender representation. This can help avoid gender-related confusion or bias in the AI model's outputs.                                                                                                                                                               |
| remove\_background    | bool   | -> When the `remove_background` parameter is enabled, the background of the generated image will be removed and replaced with a transparent background.                                                                                                                                                                                                                              |
| bg\_color             | string | -> The `bg_color` parameter allows you to set the background color of the generated image. The color value should be provided in hexadecimal format (e.g., `#FFFFFF` for white).                                                                                                                                                                                                     |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

\
Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/face-avatar-3>" %}


# Image Variation

This model has the ability to create diverse visual image variations by drawing inspiration from an input image.

For more details on this model, please refer to the [Broken mention](broken://pages/jVXXwaFgF4TUXOeUU8MR) section.

## Image variation API&#x20;

<mark style="color:green;">`POST`</mark> `/getImagetoImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name                  | Type   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                |
| --------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| app\_id               | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                                                                                                                                               |
| image                 | string | <p>-> The <code>image</code> parameter specifies the URL of an existing image that will be used as a reference for the generation process. </p><p>-> If the original image dimensions exceed 1536x1536 pixels, the image will be adjusted to fit within this size while preserving the original aspect ratio.</p>                                                                                                                          |
| prompt                | string | -> The `prompt` parameter is the textual input that guides the image generation process. This prompt serves as an artistic compass, shaping the visual output.                                                                                                                                                                                                                                                                             |
| guidance\_scale       | float  | <p>-> The <code>guidance\_scale</code> parameter determines how closely the generated image adheres to the provided prompt. Higher values result in the model following the prompt more closely, while lower values allow for more creative deviation.</p><p>-> The valid range for the <code>guidance\_scale</code> parameter is between 1 and 30.</p>                                                                                    |
| batch                 | int    | <p>-> The <code>batch</code> parameter allows you to specify the number of images to generate at once. </p><p>-> The valid range for this parameter is between 1 and 8.</p>                                                                                                                                                                                                                                                                |
| strength              | float  | <p>-> The <code>strength</code> parameter specifies the degree of transformation applied to the reference image. </p><p>-> A higher <code>strength</code> value (up to 1) results in the generated image closely following the initial reference image.</p>                                                                                                                                                                                |
| negative\_prompt      | string | -> The `negative_prompt` parameter allows you to specify content that you want the image generation model to avoid or minimize in the output. This can be useful for excluding certain visual elements or styles that you do not want to be present in the generated image.                                                                                                                                                                |
| num\_inference\_steps | int    | <p>-> The <code>num\_inference\_steps</code> parameter represents the number of denoising iterations to perform during the image generation process. Generally, more iterations can result in higher-quality images, but they also increase the time required for generation.</p><p>-> The valid range for the <code>num\_inference\_steps</code> parameter is between 1 and 50.</p>                                                       |
| celery                | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.                                                                                                                                    |
| inference\_type       | string | <p>-> The <code>inference\_type</code> parameter allows you to specify the GPU to be used for the image generation task. The supported values are:</p><ul><li><code>a10g</code></li><li><code>a100</code></li><li><code>h100</code></li></ul><p>The different GPU options provide varying levels of performance and capabilities, allowing you to choose the most suitable GPU based on your requirements and the demand for the task.</p> |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/image-variation>" %}


# Illusion Diffusion

**Illusion Diffusion Model**

Illusion Diffusion is a model built on the foundation of SD 1.5, designed to create illusionary effects within a given image. To use this model, you need to provide an input image or pattern, along with a prompt. The model will then generate a new image where the original image or pattern is concealed within the resulting artwork.

For more details on the Illusion Diffusion model, please refer to the [Broken mention](broken://pages/GgEcUUDai9Tk8AgXKxFE) section.

## Illusion Diffusion API

<mark style="color:green;">`POST`</mark> `/getImagetoImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name                  | Type   | Description                                                                                                                                                                                                                                                                                                                                                                                                                                |
| --------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| app\_id               | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                                                                                                                                               |
| image                 | string | <p>-> The <code>image</code> parameter specifies the URL of an existing image that will be used as a reference for the generation process. </p><p>-> If the original image dimensions exceed 1536x1536 pixels, the image will be adjusted to fit within this size while preserving the original aspect ratio.</p>                                                                                                                          |
| prompt                | string | -> The `prompt` parameter is the textual input that guides the image generation process. This prompt serves as an artistic compass, shaping the visual output.                                                                                                                                                                                                                                                                             |
| guidance\_scale       | float  | <p>-> The <code>guidance\_scale</code> parameter determines how closely the generated image adheres to the provided prompt. Higher values result in the model following the prompt more closely, while lower values allow for more creative deviation.</p><p>-> The valid range for the <code>guidance\_scale</code> parameter is between 1 and 30.</p>                                                                                    |
| batch                 | int    | <p>-> The <code>batch</code> parameter allows you to specify the number of images to generate at once. </p><p>-> The valid range for this parameter is between 1 and 8.</p>                                                                                                                                                                                                                                                                |
| strength              | float  | <p>-> The <code>strength</code> parameter specifies the degree of transformation applied to the reference image. </p><p>-> A higher <code>strength</code> value (up to 1) results in the generated image closely following the initial reference image.</p>                                                                                                                                                                                |
| negative\_prompt      | string | -> The `negative_prompt` parameter allows you to specify content that you want the image generation model to avoid or minimize in the output. This can be useful for excluding certain visual elements or styles that you do not want to be present in the generated image.                                                                                                                                                                |
| num\_inference\_steps | int    | <p>-> The <code>num\_inference\_steps</code> parameter represents the number of denoising iterations to perform during the image generation process. Generally, more iterations can result in higher-quality images, but they also increase the time required for generation.</p><p>-> The valid range for the <code>num\_inference\_steps</code> parameter is between 1 and 50.</p>                                                       |
| celery                | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.                                                                                                                                    |
| inference\_type       | string | <p>-> The <code>inference\_type</code> parameter allows you to specify the GPU to be used for the image generation task. The supported values are:</p><ul><li><code>a10g</code></li><li><code>a100</code></li><li><code>h100</code></li></ul><p>The different GPU options provide varying levels of performance and capabilities, allowing you to choose the most suitable GPU based on your requirements and the demand for the task.</p> |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/illusion-diffusion>" %}


# Upscaling

The Super-Resolution model, inspired by Real-ESRGAN, allows you to magnify the dimensions of any given image. This tool is excellent for upscaling photographs, illustrations, and graphics, helping you unlock greater levels of detail and clarity in your images.

For more details on the Super-Resolution model, please refer to the [Broken mention](broken://pages/VPM0zZSJGRj0eYEgmTe7) section.

## Upscale API

<mark style="color:green;">`POST`</mark> `/getImagetoImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name    | Type   | Description                                                                                                                                                                                                                                                                                                       |
| ------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| app\_id | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                      |
| image   | string | <p>-> The <code>image</code> parameter specifies the URL of an existing image that will be used as a reference for the generation process. </p><p>-> If the original image dimensions exceed 1536x1536 pixels, the image will be adjusted to fit within this size while preserving the original aspect ratio.</p> |
| scale   | int    | -> The `scale` parameter determines the degree of upscaling applied to the input image. The available values for this parameter are 2, 4, and 8.                                                                                                                                                                  |
| celery  | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.           |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/upscaling>" %}


# Background Removal

The Background Removal model offers a versatile solution for removing the background from an image. By default, it replaces the background with a transparent backdrop.

You can also customize the background by specifying the RGB values and activating the `bg_color` parameter. Additionally, you can choose to blur the background or replace it with a different image.

This model provides a flexible tool for enhancing your images by transforming and personalizing the setting to suit your vision. For more details on this model, please refer to the [Broken mention](broken://pages/N9PrN6Bz3ihq0O05mZpK) section.

## Remove Background API

<mark style="color:green;">`POST`</mark> `/getImagetoImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name      | Type   | Description                                                                                                                                                                                                                                                                                                       |
| --------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| app\_id   | string | -> Each model is uniquely characterized by its own `app_id`.                                                                                                                                                                                                                                                      |
| image     | number | <p>-> The <code>image</code> parameter specifies the URL of an existing image that will be used as a reference for the generation process. </p><p>-> If the original image dimensions exceed 1536x1536 pixels, the image will be adjusted to fit within this size while preserving the original aspect ratio.</p> |
| bg\_img   | string | -> The `bg_image` parameter allows you to specify a URL for an image that will be used as the background for the generated image.                                                                                                                                                                                 |
| bg\_color | bool   | -> To use a custom background color, set the `bg_color` parameter to `true`. This will allow you to specify the desired RGB color values for the background.                                                                                                                                                      |
| r\_color  | int    | -> The `r_color` parameter controls the strength of the red hue in the background color, with a range from 0 to 255. Higher values result in a more vibrant red, while lower values make the red more subdued.                                                                                                    |
| g\_color  | int    | -> The `g_color` parameter controls the strength of the green hue in the background color, with a range from 0 to 255. Higher values result in a more vibrant green, while lower values make the green more subdued.                                                                                              |
| b\_color  | int    | -> The `b_color` parameter controls the strength of the blue hue in the background color, with a range from 0 to 255. Higher values result in a more vibrant blue, while lower values make the blue more subdued.                                                                                                 |
| blur      | bool   | -> Enable the `blur` parameter to apply a blurring effect to the background of the generated image, creating a captivating and enchanting visual effect.                                                                                                                                                          |
| celery    | bool   | -> The `celery` parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique `task_id`. This `task_id` allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.           |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/background-removal>" %}


# Text to Speech

This innovative model has the potential to harness the capabilities of state-of-the-art technology, enabling the creation of lifelike, enthralling speech across a diverse array of languages. The more detail about this feature could be found on this [Broken mention](broken://pages/iqUAE6SI8KMlzDAeHTPS).

## Generate Speech

<mark style="color:green;">`POST`</mark> `/getAudio`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

<table><thead><tr><th>Name</th><th>Type</th><th>Description</th></tr></thead><tbody><tr><td>app_id</td><td>string</td><td>-> Each model is uniquely characterized by its own <code>app_id</code>.</td></tr><tr><td>prompt</td><td>string</td><td><p>-> The <code>prompt</code> parameter is the textual input that guides the audio generation process. This prompt serves as an artistic compass, shaping the audio output.</p><p>-> The minimum length of the prompt is 10 characters, and the maximum length is 2500 characters.</p></td></tr><tr><td>generate_audio</td><td>bool</td><td>-> Enable the <code>generate_audio</code> parameter to generate audio output in the form of speech.</td></tr><tr><td>audio_parameters</td><td>dict/map</td><td><p>-> The <code>audio_parameters</code> parameter is a dictionary that allows you to specify various audio-related settings. Here's an example of the structure:</p><pre><code>"audio_parameters": {
  "voice_id": "21m00Tcm4TlvDq8ikWAM",
  "stability": 0.5,
  "similarity_boost": 0.75,
  "style": null,
  "use_speaker_boost": true
}
</code></pre><p>You can customize the values within this dictionary to adjust the audio generation according to your preferences.</p></td></tr><tr><td>celery</td><td>bool</td><td>-> The <code>celery</code> parameter is used for queuing tasks that require extended processing time. When you enqueue a task, you receive a unique <code>task_id</code>. This <code>task_id</code> allows you to check the task's status later using the task status API, which is useful for managing and tracking long-running tasks.</td></tr></tbody></table>

The `audio_parameters` dictionary contains the following parameters:

<table><thead><tr><th width="191">Name</th><th width="166">Type</th><th>Description</th></tr></thead><tbody><tr><td>voice_id</td><td>string</td><td><p>-> The <code>voice_id</code> parameter specifies a unique identifier for the voice to be used in the audio generation process. Some  supported voice_id and their Attributes are-<br></p><ul><li><strong>VoiceID:</strong> EXAVITQu4vr4xnSDxMaL<br><strong>Name:</strong> Sarah<br><strong>Attributes:</strong> american, professional, young, female, en, entertainment_tv</li><li><strong>VoiceID:</strong> N2lVS1w4EtoT3dr4eOWO<br><strong>Name:</strong> Callum<br><strong>Attributes:</strong> en, middle_aged, male, characters</li><li><strong>VoiceID:</strong> JBFqnCBsd6RMkjVDRZzb<br><strong>Name:</strong> George<br><strong>Attributes:</strong> british, mature, middle_aged, male, en, narrative_story</li><li><strong>VoiceID:</strong> pqHfZKP75CvOlQylNhV4<br><strong>Name:</strong> Bill<br><strong>Attributes:</strong> american, crisp, old, male, en, advertisement</li><li><strong>VoiceID:</strong> NFG5qt843uXKj4pFvR7C<br><strong>Name:</strong> Adam Stone - late night radio<br><strong>Attributes:</strong> british, meditative, middle_aged, male, en, narrative_story</li><li><strong>VoiceID:</strong> XrExE9yKIg1WjnnlVkGX<br><strong>Name:</strong> Matilda<br><strong>Attributes:</strong> american, upbeat, middle_aged, female, en, informative_educational</li></ul></td></tr><tr><td>stability</td><td>float</td><td>-> The <code>stability</code> parameter controls the stability of the generated audio. Higher values (up to 1) result in more stable output, while lower values can lead to more variable output.</td></tr><tr><td>similarity_boost</td><td>float</td><td>-> The <code>similarity_boost</code> parameter adjusts the similarity boost applied to the generated audio. Higher values (up to 1) result in output that is more similar to the target, while lower values can lead to more variation.</td></tr><tr><td>style</td><td>float</td><td><p>-> The <code>style</code> parameter allows you to control the style of the generated speech. Higher values (up to 1) can result in more exaggerated or closely following the given voice style, but may also lead to increased instability in the generated speech.</p><p>-> Setting this parameter to 0.0 (the default) will greatly increase the generation speed.</p></td></tr><tr><td>use_speaker_boost</td><td>bool</td><td>-> When enabled, the <code>use_speaker_boost</code> parameter will boost the similarity of the synthesized speech to the selected voice, at the cost of some generation speed. This option makes the model try to generate speech that more closely aligns with the chosen voice.</td></tr></tbody></table>

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="400" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/text-to-speech>" %}


# Text to Speech

Converts text into lifelike audio using Google's Gemini TTS model. Supports multilingual synthesis, a wide selection of distinct voices, and optional style prompting to control tone and delivery.

## Generate Speech

**POST** `/api/v1/studio/synthesizeSpeech`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**<br>

| Parameter       | Type     | Required | Description                                                                                                                                                                                                                                   |
| --------------- | -------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `text`          | `string` | **Yes**  | The input text to be converted into speech. Must be a non-empty string. This is the raw content that will be spoken aloud in the generated audio.                                                                                             |
| `voice`         | `string` | **Yes**  | The name of the voice persona to use for synthesis. Each voice has a distinct character and tone. Must be one of the supported voices (see `GET /api/v1/studio/speech/voices`). Examples: `"Zephyr"`, `"Puck"`, `"Aoede"`, `"Kore"`.          |
| `model`         | `string` | No       | The TTS model to use for generation. Defaults to `"gemini-2.5-flash-tts"`                                                                                                                                                                     |
| `language_code` | `string` | No       | BCP-47 language code to control the language of the synthesized speech (e.g., `"en-US"`, `"fr-FR"`, `"hi-IN"`). If omitted, the model infers the language from the input text. See `GET /api/v1/studio/speech/languages` for supported codes. |
| `style_prompt`  | `string` | No       | A natural-language instruction that shapes the speaking style, tone, or emotion of the output. Examples: `"Speak in a calm, soothing tone"`, `"Sound excited and energetic"`, `"Read like a news anchor"`.                                    |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "url": "",
  "mime_type": "",
  "referenceID": ""
}
```

{% endtab %}

{% tab title="403" %}

```json
{
  "message": "Forbidden: Insufficient credit"
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "message": "Internal Server Error"
}
```

{% endtab %}

{% tab title="401" %}

```json
{
  "message": "Missing authorization token"
}
```

{% endtab %}

{% tab title="400" %}

```json
{
  "message": "text is required and cannot be empty."
}
```

{% endtab %}
{% endtabs %}

| Field         | Type     | Description                                                                           |
| ------------- | -------- | ------------------------------------------------------------------------------------- |
| `url`         | `string` | A publicly accessible URL to the generated audio file.                                |
| `mime_type`   | `string` | The MIME type of the audio file (e.g., `"audio/mpeg"`, `"audio/wav"`).                |
| `referenceID` | `string` | A unique identifier for this synthesis request, used for tracking and audit purposes. |

***

#### Error Responses

| Status | Condition                                      |
| ------ | ---------------------------------------------- |
| `400`  | `text` is missing or empty                     |
| `400`  | `voice` is missing                             |
| `400`  | `voice` is not in the list of supported voices |
| `500`  | Internal server error                          |

***

#### Related Endpoints

* `GET /api/v1/studio/speech/voices` — Returns all supported voice names
* `GET /api/v1/studio/speech/languages` — Returns all supported language codes

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/text-to-speech-copy-1>" %}


# Task Status

Use this API endpoint to check the progress of a scheduled task through Celery. All you need is the `task_id`, which you receive when you schedule the task. This allows you to monitor the status and progress of your long-running tasks.

## Task Status API

<mark style="color:green;">`POST`</mark> `/getImage`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name  | Type   | Description                                                                                                                                                                                                                                                 |
| ----- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| id    | string | -> The `id` parameter is used to provide the unique identifier of the scheduled task you want to check the status for.                                                                                                                                      |
| refID | string | -> The `reference_id` parameter is an optional field that allows you to provide a reference ID to identify the request, if required. If you did not specify a `reference_id` when passing the input parameters, there is no need to provide this parameter. |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "time_required": "",
  "error": "",
  "error_data": "",
  "input": "",
  "output": "",
  "app_id": "",
  "task_id": "",
  "status": ""
}
```

{% endtab %}
{% endtabs %}

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/get-status>" %}


# StreamChat

Bot API allows you to try out different Large Language Models (LLMs) from providers like Mistral, OpenAI, and Claude.

For more details on the supported LLM models and their capabilities, please refer to the [Broken mention](broken://pages/zF6PdCmIL9GuDtdhKuvb).

## Stream chat with LLM

<mark style="color:green;">`POST`</mark> `/stream_chat`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

<table><thead><tr><th width="166">Name</th><th width="162">Type</th><th>Description</th></tr></thead><tbody><tr><td>llm</td><td>string</td><td><p>-> The <code>llm</code> parameter specifies the type of Large Language Model (LLM) to be used. The supported values are:</p><ul><li><code>ClaudeAI</code></li><li><code>OpenAI</code></li><li><code>GeminiAI</code></li><li><code>OpenRouterAI</code></li></ul><p></p></td></tr><tr><td>llm_model</td><td>string</td><td><p>-> The <code>llm_model</code> parameter specifies the name of the Large Language Model (LLM) to be used. The supported values for this parameter depend on the <code>llm</code> you have selected.</p><p>For example, if the <code>llm</code> is <code>OpenAI</code>, the supported <code>llm_model</code> values include:</p><ul><li><code>"gpt-4.1-mini-2025-04-14"</code></li><li><code>"gpt-4.1-2025-04-14"</code></li></ul><p>Similarly, if the <code>llm</code> is <code>ClaudeAI</code>, the supported <code>llm_model</code> values include:</p><ul><li><code>"claude-3-7-sonnet-latest"</code></li></ul><p>Ensure that the <code>llm_model</code> value you provide is compatible with the selected <code>model_type</code>.</p><p></p><p>-> Supported LLMs and their Models are : </p><p></p><p>GeminiAI</p><ol><li>gemini-2.5-pro </li><li>gemini-2.5-flash</li></ol><p></p><p>ClaudeAI</p><ol><li>claude-sonnet-4-20250514</li><li>claude-opus-4-1-20250805</li></ol><p></p><p>OpenAI</p><ol><li>gpt-4.1-mini-2025-04-14</li><li>gpt-4.1-nano-2025-04-14</li><li>gpt-4o-mini </li><li>gpt-4.1-2025-04-14</li><li>o1</li></ol><p>(Note: Below models for OpenAI LLM must be configured with <code>temperature=1</code> ,while runninng this model)</p><ol><li>o3-mini</li><li>o3</li><li>o4-mini-2025-04-16  </li><li>gpt-5-2025-08-07</li><li>gpt-5-mini-2025-08-07</li><li>gpt-5-nano-2025-08-07</li></ol><p></p><p>OpenRouterAI</p><ol><li>x-ai/grok-4</li><li>x-ai/grok-3-mini-beta</li><li>x-ai/grok-3-beta</li><li>perplexity/sonar-pro</li><li>perplexity/sonar-reasoning-pro</li><li>perplexity/sonar-reasoning</li><li>perplexity/sonar-deep-research</li><li>deepseek/deepseek-chat</li><li>deepseek/deepseek-r1</li></ol></td></tr><tr><td>history</td><td>List of Dictionary (Python) or map (C++)</td><td><p>-> The <code>history</code> parameter allows you to provide the previous chat history and the last user message. The history should follow a specific pattern:</p><ol><li>The history should start with a user message.</li><li>After each user message, there should be a response from the assistant.</li><li>The last message in the history should be a user message or query.</li></ol><p>-> The history should be provided as a list of dictionaries, where each dictionary represents a message. The dictionary should have the following structure:</p><pre><code> {
    "role": "user",
    "content": {
      "text": "please give 50 word story in English",
      "image_data": [
         {
           "url": "string",
           "details": "low"
         }
       ]
     }
  }

</code></pre><p>The <code>image\_data</code> parameter is a list that can contain up to 5 image URLs. The <code>details</code> parameter is optional and can be used to specify the level of detail required for the image analysis (only for OpenAI models).</p><p>Here's an example of the <code>history</code> parameter:</p><pre><code>"history": \[
{
"role": "user",
"content": {
"text": "hey",
"image\_data": \[]
}
},
{
"role": "assistant",
"content": {
"text": "Hello! How can I assist you today?",
"image\_data": \[]
}
},
{
"role": "user",
"content": {
"text": "Please analyze this image",
"image\_data": \[
{
"url": "<https://res.cloudinary.com/qolaba/image/upload/v1695690455/kxug1tmiolt1dtsvv5br.jpg>",
"details": "low"
}
]
}
}
]

</code></pre></td></tr><tr><td>temperature</td><td>float</td><td><p>-> The <code>temperature</code> parameter accepts a float value between 0 and 1. This parameter helps control the level of determinism in the output from the Large Language Model (LLM).</p><p>-> Higher temperature values (closer to 1) will result in more diverse and creative output, while lower values (closer to 0) will lead to more deterministic and conservative responses.</p></td></tr><tr><td>image\_analyze</td><td>bool</td><td>-> If you are passing image URLs and want the model to analyze the images, set the <code>image\_analyze</code> parameter to <code>true</code>.</td></tr><tr><td>enable\_tool</td><td>bool</td><td><p>-> To use the tools supported by the Chat API, enable the <code>enable\_tool</code> parameter. The Chat API currently supports two tools:</p><ol><li>Vector Search</li><li>Internet Search</li></ol><p>-> After enabling the <code>enable\_tool</code> parameter, you can provide the details of the tool you want to use in the <code>tools</code> parameter.<br></p></td></tr><tr><td>system\_msg</td><td>string</td><td>-> The <code>system\_msg</code> parameter allows you to set a system message for the Large Language Model (LLM). This message can be used to provide context or instructions to the model, which can influence the tone and behavior of the generated responses.</td></tr><tr><td>tools</td><td>dictionary or map</td><td><p></p><p>The <code>tools</code> parameter allows you to configure the capabilities available to the assistant during a conversation. It defines which tools are enabled, how much past context to consider, and specifies settings for embeddings, PDFs, and image generation. This is useful for tailoring the assistant’s functionality to your use case.</p><p><strong>🔹 Structure</strong></p><pre class="language-json"><code class="lang-json">jsonCopyEdit"tools": {
"tool\_list": {
"image\_generation": false,
"image\_generation1": false,
"image\_editing": false,
"search\_doc": false,
"internet\_search": false,
"python\_code\_execution\_tool": false,
"csv\_analysis": false
},
"number\_of\_context": 3,
"pdf\_references": \[],
"embedding\_model": \[
"text-embedding-3-large"
],
"image\_generation\_parameters": {}
} </code></pre><p><strong>🔹 Field Descriptions</strong></p><ul><li><p><strong>tool\_list</strong>: An object with boolean flags to enable or disable specific tools. Setting a value to <code>true</code> activates that feature.</p><ul><li><code>image\_generation</code>: Enables image creation from text prompts.</li><li><code>image\_generation1</code>: Optional alternate switch for image generation (if supported).</li><li><code>image\_editing</code>: Enables uploading and editing images.</li><li><code>search\_doc</code>: Allows document-based searching.</li><li><code>internet\_search</code>: Enables real-time web searches.</li><li><code>python\_code\_execution\_tool</code>: Allows execution of Python code within a safe sandbox.</li><li><code>csv\_analysis</code>: Enables processing and analysis of uploaded CSV files.</li></ul></li><li><strong>number\_of\_context</strong>: An integer defining how many previous user-assistant interactions to retain in memory. Higher values help preserve conversation flow.</li><li><strong>pdf\_references</strong>: A list of PDFs that the assistant can reference for contextual information should be stored in vector store database. Each entry includes an unique id which comes from the output of vector store API. If you're adding more than one document using <code>pdf\_references</code>, make sure you add the same number of entries in the <code>embedding\_model</code> array.</li><li><strong>embedding\_model</strong>: we support  <code>"text-embedding-3-large"</code> as  the embedding model.</li><li><strong>image\_generation\_parameters</strong>: An object to define settings like size or style for image generation. Leave it empty for default behavior.</li></ul></td></tr><tr><td>token</td><td>string</td><td>Authentication token needed for the request.</td></tr><tr><td>orgID</td><td>string </td><td>Identifies the organization associated with the request.</td></tr><tr><td>function\_call\_list</td><td>string</td><td>Lists the functions called during the process.</td></tr><tr><td>systemId </td><td>string</td><td>Specifies the system identifier linked to the request.</td></tr><tr><td>last\_user\_query</td><td>string</td><td>Provides the last user query made in the chat.</td></tr></tbody></table>

Example of Input body parameters :&#x20;

```
data '{
  "llm": "GeminiAI",
  "llm_model": "gemini-2.5-flash",
  "history": [
    {
      "role": "user",
      "content": {
        "text": "please give 50 word story in English",
        "image_data": [
          {
            "url": "string",
            "details": "low"
          }
        ]
      }
    }
  ],
  "temperature": 0.7,
  "image_analyze": false,
  "enable_tool": false,
  "system_msg": "You are a helpful assistant.",
  "tools": {
    "tool_list": {
      "image_generation": false,
      "image_generation1": false,
      "image_editing": false,
      "search_doc": false,
      "internet_search": false,
      "python_code_execution_tool": false,
      "csv_analysis": false
    },
    "number_of_context": 3,
      "pdf_references": [
      "output vectorstore1",
      "output vectorstore2"
    ],
    "embedding_model": [
      "text-embedding-3-large",
      "text-embedding-3-large"
    ],
    "image_generation_parameters": {}
  },
  "token": "123",
  "orgID": "string",
  "function_call_list": [],
  "systemId": "string",
  "last_user_query": "string"
}'
```

After passing the necessary parameters and executing the Chat API, you will receive a stream response. A successful response will look similar to the following:

**Response**

{% tabs %}
{% tab title="200" %}

<pre class="language-json"><code class="lang-json"><strong>{
</strong>  "output": null, 
  "error": null, 
  "error_data": null
}
</code></pre>

{% endtab %}

{% tab title="500" %}

```json
{
  "output": null, 
  "error": string, 
  "error_data": string
}
```

{% endtab %}
{% endtabs %}

The Chat API response is a streaming response, which means you will receive the output in chunks as the model generates the response. The response will continue to stream until the generation is complete.

The response will contain the following elements:

* `output`: This object contains the generated text output from the model.

The text output from the LLM can be obtained from the `output` parameter. When the response is complete, the final chunk will contain a `null` value in the `output` parameter, indicating the end of the stream.

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/streamchat>" %}


# OpenAI-Compatible Chat

OpenAI-compatible chat completions endpoint. Supports both standard (JSON) and streaming (SSE) responses. Rate-limited per API key.

**Base URL:**

{% embed url="<https://api.platform.qolaba.ai>" %}

**Supported endpoints**

| `GET`  | `/v1/chat/models`          | List all available models                              |
| ------ | -------------------------- | ------------------------------------------------------ |
| `GET`  | `/v1/chat/models/{model}`  | Retrieve a specific model                              |
| `POST` | `/api/v1/chat/completions` | Create chat completions (streaming, tools, multimodal) |

## Chat Completions

**POST** `/api/v1/chat/completions`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Parameter           | Type                 | Required | Description                                                                                                                                                                                                               |
| ------------------- | -------------------- | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `model`             | `string`             | **Yes**  | The model ID to use for completion. Use `GET /api/v1/chat/models` to list available models.                                                                                                                               |
| `messages`          | `array`              | **Yes**  | Ordered list of messages forming the conversation. Each object must have a `role` (`"system"`, `"user"`, or `"assistant"`) and a `content` field (string or structured content array). Must contain at least one message. |
| `stream`            | `boolean`            | No       | If `true`, the response is streamed as Server-Sent Events (SSE). Each event is a `data: {...}` line; the stream ends with `data: [DONE]`. Defaults to `false`.                                                            |
| `temperature`       | `number`             | No       | Sampling temperature between `0` and `2`. Higher values produce more random output; lower values are more deterministic. Defaults to `1`.                                                                                 |
| `max_tokens`        | `integer`            | No       | Maximum number of tokens to generate in the response. Must be a positive integer. If omitted, the model's default limit applies.                                                                                          |
| `top_p`             | `number`             | No       | Nucleus sampling threshold between `0` and `1`. The model considers only the tokens comprising the top `top_p` probability mass. An alternative to `temperature`; avoid adjusting both simultaneously.                    |
| `frequency_penalty` | `number`             | No       | Penalizes tokens based on how frequently they have appeared so far. Positive values reduce repetition. Typically between `-2.0` and `2.0`.                                                                                |
| `presence_penalty`  | `number`             | No       | Penalizes tokens that have appeared at all in the conversation so far. Positive values encourage the model to introduce new topics. Typically between `-2.0` and `2.0`.                                                   |
| `stop`              | `string \| string[]` | No       | One or up to **4** sequences at which the model will stop generating further tokens. The stop sequence itself is not included in the output.                                                                              |
| `tool_choice`       | `string \| object`   | No       | Controls which tool (if any) the model calls. Accepts `"none"`, `"auto"`.                                                                                                                                                 |
| `stream_options`    | `object`             | No       | Additional options for streaming. Set `{ "include_usage": true }` to receive a final SSE chunk containing token usage data. Automatically enabled when `stream: true`.                                                    |

#### Request

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "What is the capital of France?"
    }
  ],
  "stream": false
}
```

#### Response (non-streaming)

{% tabs %}
{% tab title="200" %}

```json
{
  "id": "chatcmpl-4dd9fded095b4d86bceed1e91042a",
  "object": "chat.completion",
  "created": 1774095579,
  "model": "google/gemini-2.5-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Paris."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "total_credit": 0.25,
    "total_cost": 0.0012441,
    "breakdown": {
      "model": {
        "model": "google/gemini-2.5-flash",
        "credit": 0.25,
        "cost": 0.001244
      }
    }
  }
}
```

{% endtab %}
{% endtabs %}

#### Error Responses

| Status | Condition                                     |
| ------ | --------------------------------------------- |
| `400`  | `model` is missing                            |
| `400`  | `messages` is missing, not an array, or empty |
| `400`  | `temperature` is outside `0–2`                |
| `400`  | `top_p` is outside `0–1`                      |
| `400`  | `max_tokens` is not a positive integer        |
| `400`  | `stop` array contains more than 4 sequences   |
| `429`  | Rate limit exceeded                           |
| `504`  | Request timed out (3-minute upstream limit)   |
| `502`  | Upstream provider error or invalid response   |
| `500`  | Internal server error                         |

***

#### Related Endpoints

* `GET /api/v1/chat/models` — Returns the list of available models<br>

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/chat-copy>" %}


# Tool Calls

The API supports OpenAI-compatible function calling. Define your functions in the request; the model decides when to call them, and the server returns the call details for you to execute on the client

### Step 1 — Define tools and get a tool call

```json
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: process.env.API_KEY,
  baseURL: 'http://localhost:9000/v1',
});

const completion = await client.chat.completions.create({
  model: 'openai/gpt-4o',
  messages: [
    { role: 'user', content: 'What is the weather in Tokyo?' }
  ],
  tools: [
    {
      type: 'function',
      function: {
        name: 'get_weather',
        description: 'Get the current weather for a city',
        parameters: {
          type: 'object',
          properties: {
            location: {
              type: 'string',
              description: 'City name, e.g. Tokyo, Japan'
            },
            unit: {
              type: 'string',
              enum: ['celsius', 'fahrenheit'],
              description: 'Temperature unit'
            }
          },
          required: ['location']
        }
      }
    }
  ],
  tool_choice: 'auto',
  built_in_tools: 'none'  // disable server tools so only yours are offered
});

console.log(completion.choices[0].message.tool_calls);
```

**Response — model decides to call your function**

```json
{
  "id": "chatcmpl-xyz",
  "object": "chat.completion",
  "created": 1740700000,
  "model": "openai/gpt-4o",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_abc123",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"location\": \"Tokyo, Japan\", \"unit\": \"celsius\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ]
}
```

### Step 2 — Execute the function and send the result back

Parse `arguments`, call your function, then continue the conversation with a `tool` message:

```json
// Parse arguments and call your function
const toolCall = completion.choices[0].message.tool_calls[0];
const args = JSON.parse(toolCall.function.arguments);

// Your own function execution
const weatherResult = await myGetWeather(args.location, args.unit);

// Continue conversation with the tool result
const followUp = await client.chat.completions.create({
  model: 'openai/gpt-4o',
  messages: [
    { role: 'user', content: 'What is the weather in Tokyo?' },
    completion.choices[0].message,  // include the assistant's tool_calls message
    {
      role: 'tool',
      tool_call_id: toolCall.id,
      content: JSON.stringify(weatherResult)
    }
  ],
  tools: [ /* same tool definitions */ ],
  built_in_tools: 'none'
});

console.log(followUp.choices[0].message.content);
// "The current weather in Tokyo is 12°C with partly cloudy skies."
```

### Controlling tool selection

| `tool_choice` value                                    | Behaviour                                                |
| ------------------------------------------------------ | -------------------------------------------------------- |
| `"auto"` (default)                                     | Model decides whether to call a tool or respond directly |
| `"none"`                                               | Model must not call any tool                             |
| `"required"`                                           | Model must call at least one tool                        |
| `{ "type": "function", "function": { "name": "fn" } }` | Model must call this exact function                      |

```
// Force a specific function
"tool_choice": { "type": "function", "function": { "name": "get_weather" } }
```

### Multi-tool call in one turn

The model can call multiple tools at once:

```json
{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_001",
            "type": "function",
            "function": { "name": "get_weather", "arguments": "{\"location\":\"Tokyo\"}" }
          },
          {
            "id": "call_002",
            "type": "function",
            "function": { "name": "get_weather", "arguments": "{\"location\":\"London\"}" }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ]
}
```

Send all tool results back in one follow-up request, one `tool` message per call:

```json
{
  "messages": [
    { "role": "user", "content": "Compare weather in Tokyo and London." },
    { "role": "assistant", "content": null, "tool_calls": [ ...both calls... ] },
    { "role": "tool", "tool_call_id": "call_001", "content": "{\"temp\": 12, \"condition\": \"Cloudy\"}" },
    { "role": "tool", "tool_call_id": "call_002", "content": "{\"temp\": 8, \"condition\": \"Rainy\"}" }
  ]
}
```


# Built-in Tools

The server's agent has 4 built-in tools that run automatically on the server side — no client execution required. The model calls them internally and the result is integrated into the final response.

### Available built-in tools

| Tool               | Trigger                                         | What it does                                           |
| ------------------ | ----------------------------------------------- | ------------------------------------------------------ |
| `image_generation` | "Generate an image of a sunset over the ocean"  | Generates images; URL embedded in response as markdown |
| `internet_search`  | "Latest news about AI", "What is today's date?" | Web search via Perplexity AI with citations            |
| `file_search`      | "Summarize the uploaded document"               | Searches vector stores from uploaded files             |
| `url_content`      | "Summarize <https://example.com/article>"       | Fetches and reads content from URLs                    |

### image generation

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "Generate an image of a futuristic city at sunset." }
  ]
}
```

The generated image URL is embedded as a markdown image inside `content`. Image cost is reported in `usage.tool_cost_usd`:

```json
{
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "I've generated the image of a futuristic city at sunset.\n\n![Generated Image](https://cdn.example.com/gen/abc123.png)"
      },
      "finish_reason": "stop"
    }
  ]
}
```

### File Search

The `file_search` tool lets the model search documents in pre-indexed file stores. Pass the store names in the request:

```json
{  
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "Summarise the Q4 report" }
  ],
  "file_search_store_names": ["fileSearchStores/abc123"]
}  
```

The model will automatically call `file_search` with the provided store names when the query is relevant to the documents.

Store names are in the format `fileSearchStores/{storeId}`. You can pass multiple stores:<br>

```
"file_search_store_names": [
  "fileSearchStores/abc123",
  "fileSearchStores/def456"
]
```

### Internet Search

The `internet_search` tool queries the web via **Perplexity AI** and returns a synthesized, up-to-date answer with source citations. Use it for any question that requires current information beyond the model's training cutoff.

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "What are the latest AI news today?"
    }
  ],
  "stream": false
}
```

### URL Content

The `url_content` tool fetches and processes the content of one or more URLs. It uses **Gemini 2.5 Flash** with Google's URL Context capability to retrieve, read, and reason over live web pages.

Use it when the user provides a specific link and asks questions about it — summarisation, extraction, comparison, Q\&A over a page, etc.

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Summarize this page and also team details: https://www.qolaba.ai/about-us"
    }
  ],
  "stream": false
}
```


# Image Generation

\
The `image_generation` tool is a **built-in tool** automatically available on all `/v1/chat/completions` requests. It uses **Google Gemini** (`gemini-3.1-flash-image-preview`) under the hood. You don't configure it — just describe what you want in natural language and the model decides when to invoke it.

#### How It Works <a href="#how-it-works" id="how-it-works"></a>

1. You send a normal chat request with an image-related prompt
2. The LLM detects it needs to generate an image and calls the `image_generation` tool internally
3. The tool generates the image and uploads it to GCP CDN
4. The response comes back as a standard chat completion with a markdown image link embedded in the `content`

**Resolutions**

| Value  | Description                        |
| ------ | ---------------------------------- |
| `0.5K` | 512px — quick previews, thumbnails |
| `1K`   | Default — standard quality         |
| `2K`   | High quality                       |
| `4K`   | Maximum detail, upscaling          |

**Aspect Ratios**

| Value         | Best For                                          |
| ------------- | ------------------------------------------------- |
| `1:1`         | Social media posts, profile pictures              |
| `16:9`        | Landscape, YouTube thumbnails, desktop wallpapers |
| `9:16`        | Stories, Reels, TikTok                            |
| `4:3`         | Presentations, traditional photos                 |
| `3:2`         | Photography, prints                               |
| `21:9`        | Cinematic, ultrawide banners                      |
| `4:5` / `5:4` | Instagram portrait/landscape                      |
| `1:4` / `4:1` | Tall/wide banners                                 |

### Use Cases & Examples

#### 1. Simple Image Generation

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "Generate an image of a futuristic city at sunset with flying cars" }
  ]
}

```

#### 2. Specific Aspect Ratio + Resolution

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Generate a 16:9 high quality landscape image of snow-capped mountains with a lake reflection"
    }
  ]
}

```

The model automatically infers `resolution: 2K` and `aspect_ratio: 16:9` from the prompt.

#### **3. Multiple Image Variations**

Returns multiple `![Generated Image](url)` links in the content, cost split per image.

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "Generate 3 variations of a minimalist logo for a coffee shop" }
  ]
}

```

#### 4. Portrait / Story Format

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "Create a 9:16 portrait illustration of a woman reading a book in a cozy library" }
  ]
}

```

#### 5. High-Detail / Infographic (Thinking Mode)

Best for text-heavy images, menus, charts, complex multi-element compositions:

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "Create a detailed restaurant menu layout with sections for appetizers, mains, and desserts with prices" }
  ]
}

```

#### 6. Branded / Product Marketing

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Create a 4:5 product photo of a luxury perfume bottle on a marble surface with soft lighting, suitable for Instagram"
    }
  ]
}

```

### Tips & Best Practices

| Tip                     | Details                                                                            |
| ----------------------- | ---------------------------------------------------------------------------------- |
| **Be specific**         | Include style (photorealistic, cartoon, watercolor), lighting, mood, color palette |
| **Mention ratio**       | Say "16:9 landscape" or "9:16 vertical" — the model picks up the hint              |
| **Multiple variations** | Say "3 variations" or "4 different versions" to get `count > 1`                    |
| **Text in images**      | Mention "clear readable text" — model enables thinking mode automatically          |
| **Cost awareness**      | 4K images and multiple images multiply cost; 0.5K is cheapest for previews         |

***

### Limitations

* Max `count` per request: no hard cap but cost scales linearly
* `use_search_grounding` adds latency and extra cost per query
* Generated images are hosted on `cdn.qolaba.app` — URLs are permanent CDN links


# Internet Search

The `internet_search` tool is a **built-in tool** automatically available on all `/v1/chat/completions` requests. It uses **Perplexity AI Sonar** under the hood to search the live web and return a synthesized answer with citations. You don't configure it — the model invokes it automatically when the prompt requires real-time or current information.

### How It Works

1. You send a normal chat request with a question that needs current data
2. The LLM detects it needs fresh web data and calls `internet_search` internally
3. The tool queries **Perplexity Sonar** (`temperature: 0.1` for accuracy) with your query
4. Perplexity returns a synthesized answer + citation URLs
5. The model wraps the answer into a standard chat completion response

#### Streaming Behavior

Unlike image generation (which delivers the result at the end), internet search **integrates naturally with streaming**. The stream shows:

1. A `tool_calls` chunk announcing `internet_search` was invoked
2. Immediately followed by content chunks streaming the answer word-by-word
3. A final `[DONE]` chunk
4. A trailing `usage` data chunk (if `stream_options.include_usage: true`)

```json
data: {"delta":{"tool_calls":[{"function":{"name":"internet_search",...}}]}}
data: {"delta":{"content":"As"}}
data: {"delta":{"content":" of"}}
...
data: [DONE]
data: {"usage":{...}}

```

### Use Cases & Examples

#### 1. Breaking News / Current Events

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "What are the latest AI news today?" }
  ]
}
```

#### 2. Live Price / Market Data

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "What is the current price of Bitcoin?" }
  ]
}

```

#### 3. Combining Search + Reasoning&#xD;&#x20;&#xD;

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Search for the latest GPT-5 benchmarks and summarize how it compares to Claude Sonnet"
    }
  ]
}

```

#### 4. Technical / Developer Queries&#xD;&#x20;&#xD;

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    { "role": "user", "content": "What are the breaking changes in Bun 2.0?" }
  ]
}

```

### Tips & Best Practices

| Tip                          | Details                                                                           |
| ---------------------------- | --------------------------------------------------------------------------------- |
| **Ask naturally**            | No special syntax needed — the model detects when to search automatically         |
| **Be specific**              | "Latest React 19 release notes" works better than "React news"                    |
| **Date context helps**       | Adding "as of today" or "current" reinforces real-time lookup                     |
| **Streaming for UX**         | Use `stream: true` — answers start flowing within \~1s while search runs          |
| **Cost vs knowledge cutoff** | For questions within the model's training data, search won't trigger (saves cost) |
| **Citations in output**      | The model includes source links inline — these are real URLs from Perplexity      |

***

### Limitations

* Powered by `perplexity/sonar` — accuracy depends on Perplexity's index freshness
* No control over `numResults` from the request body — handled internally (default 5 citations)
* Does not return raw search results, only the synthesized Perplexity answer
* Adds latency (\~1–3s) compared to non-search requests
* `temperature` is fixed at `0.1` internally for factual accuracy — not overridable


# URL Content

The `url_content` tool is a **built-in tool** automatically available on all `/v1/chat/completions` requests. It uses **Google Gemini 2.5 Flash** with the native `urlContext` tool to fetch, render, and analyze live web pages. The model receives the full rendered page content and answers your question about it.

#### How It Works

1. You include one or more URLs anywhere in your message
2. The LLM detects URLs and calls `url_content` internally with those URLs + your question
3. Gemini 2.5 Flash fetches and processes the full page content using Google's `urlContext` tool
4. Gemini returns an answer grounded in the actual page content
5. The answer is returned as a standard chat completion

#### Use Cases & Examples

#### 1. Summarize a Page

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Summarize this page and also team details: https://www.qolaba.ai/about-us"
    }
  ]
}

```

#### 2. Extract Specific Information

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "What are the pricing plans listed on https://www.qolaba.ai/pricing?"
    }
  ]
}
```

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Compare these two pages and tell me the key differences:\nhttps://openai.com\nhttps://anthropic.com"
    }
  ]
}

```

#### 4. Extract Structured Data from a Page

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "List all team members and their roles from https://www.qolaba.ai/about-us as a JSON array"
    }
  ]
}

```

#### 5. Competitor Analysis (3+ URLs)

```json
{
  "model": "google/gemini-2.5-flash",
  "messages": [
    {
      "role": "user",
      "content": "Compare the features and positioning of these three AI platforms:\nhttps://openai.com\nhttps://anthropic.com\nhttps://www.qolaba.ai"
    }
  ]
}

```

### Tips & Best Practices

| Tip                    | Details                                                                          |
| ---------------------- | -------------------------------------------------------------------------------- |
| **Just paste the URL** | No special syntax — include the URL anywhere in the message                      |
| **Be specific**        | "Extract all product names" works better than "tell me about this page"          |
| **Multiple URLs**      | Put each URL on its own line for clarity — all are fetched together              |
| **JS-heavy pages**     | Works on SPAs and dynamic pages — Gemini's urlContext renders JavaScript         |
| **Large pages**        | Cost scales with page size but remains low (Gemini 2.5 Flash is cheap per token) |
| **Ask for format**     | "Return as JSON", "Return as a table", "List as bullet points" all work          |

***

### Limitations

* **Private/authenticated pages** — cannot access pages behind login walls or paywalls
* **PDFs / binary files** — works best on HTML pages; PDF support depends on Google's urlContext
* **Rate-limited sites** — some sites block scraping; Gemini makes a best-effort attempt
* **Very large pages** — extremely long pages may be truncated by token limits
* **No config from request** — `urls` and `prompt` are extracted by the model; you cannot pass them directly in the request body


# File Search

The `file_search` tool is a **built-in tool** for searching through your uploaded documents. Unlike `internet_search` (live web) and `url_content` (public URLs), this tool searches **your private uploaded files** stored in a **File Search Store**. It uses **Google Gemini 2.5 Flash** with native `fileSearch` grounding to retrieve and answer from document content.

### How It Works

1. You (or your system) upload documents to a **File Search Store** (via a separate upload API)
2. You pass the store name(s) in the `file_search_store_names` field of your request
3. The LLM receives the store names via an injected system message and calls the `file_search` tool automatically
4. Gemini 2.5 Flash searches the store using semantic/vector search and returns grounded answers with citations
5. The answer is returned as a standard chat completion

#### Key Difference from Other Tools

| Tool              | Trigger                                     | Source                              |
| ----------------- | ------------------------------------------- | ----------------------------------- |
| `internet_search` | Auto (from prompt)                          | Live public web via Perplexity      |
| `url_content`     | Auto (URL in message)                       | Specific public URL via Gemini      |
| `file_search`     | Auto + `file_search_store_names` in request | **Your private uploaded documents** |

#### Use Cases & Examples

#### 1. Summarize an Uploaded Document

```json
{
  "model": "google/gemini-2.5-flash",
  "file_search_store_names": ["fileSearchStores/defaultstore-31ihhh9rmyef"],
  "messages": [
    { "role": "user", "content": "Summarize the uploaded document" }
  ]
}
```

```json
{
  "model": "google/gemini-2.5-flash",
  "file_search_store_names": [
    "fileSearchStores/q1-reports",
    "fileSearchStores/q2-reports"
  ],
  "messages": [
    { "role": "user", "content": "Compare Q1 and Q2 revenue figures from the reports" }
  ]
}
```

#### 4. Contract / Legal Document Analysis

```json
{
  "model": "google/gemini-2.5-flash",
  "file_search_store_names": ["fileSearchStores/contracts-store"],
  "messages": [
    { "role": "user", "content": "What are the termination clauses in the uploaded contract?" }
  ]
}
```

#### 5. With a System Prompt

```json
{
  "model": "google/gemini-2.5-flash",
  "file_search_store_names": ["fileSearchStores/support-docs"],
  "messages": [
    {
      "role": "system",
      "content": "You are a helpful support agent. Only answer based on the uploaded documentation."
    },
    { "role": "user", "content": "How do I upgrade my plan?" }
  ]
}

```

#### Tips & Best Practices

| Tip                            | Details                                                                                 |
| ------------------------------ | --------------------------------------------------------------------------------------- |
| **Always pass store names**    | `file_search_store_names` must be in every request — the tool won't activate without it |
| **Pass on every turn**         | In multi-turn conversations, include the field on each request                          |
| **Multiple stores**            | Pass an array to search across multiple document collections at once                    |
| **Any model works**            | Store names are injected as a system message — model-agnostic                           |
| **Be specific**                | "What is the refund policy?" retrieves more precisely than "Tell me about the document" |
| **Combine with system prompt** | Use a system prompt to constrain the model to only answer from the documents            |

***

### Limitations

* **Requires pre-uploaded documents** — you must upload files to a store before querying (separate upload flow)
* **Store names are static** — if a store is deleted or renamed, queries will fail gracefully (returns an error message without crashing)
* **No raw chunk retrieval** — you get a synthesized answer, not a list of raw document excerpts
* **No filtering by file** — all files in the store are searched together; you can't target a specific file within a store
* **`file_search_store_names` is not a standard OpenAI field** — OpenAI SDK clients must use `extra_body` (Python) or `// @ts-ignore` (TypeScript) to pass it


# File Upload

Upload a file by URL to create a file search store. The returned store\_name can be used in chat completions via file\_search\_store\_names.

#### RAG File Upload

Uploads a file by URL to the AI backend for retrieval-augmented generation.

**Endpoint:** `POST /api/v1/files`

**Auth:** Required (`X-API-Key` or `Authorization: Bearer`)\
**Timeout:** 5 minutes

**Request body:**

| Parameter    | Type   | Required | Description                      |
| ------------ | ------ | -------- | -------------------------------- |
| `url`        | string | ✅        | Public URL to the file           |
| `store_name` | string | —        | Custom name for the vector store |

**Allowed file types:**

| Extension | MIME Type                                                                 |
| --------- | ------------------------------------------------------------------------- |
| `.pdf`    | `application/pdf`                                                         |
| `.txt`    | `text/plain`                                                              |
| `.doc`    | `application/msword`                                                      |
| `.docx`   | `application/vnd.openxmlformats-officedocument.wordprocessingml.document` |

#### Request:

```bash
curl -X POST https://qolaba-server-b2b.up.railway.app/api/v1/files \
  -H "Content-Type: application/json" \
  -H "x-api-key: YOUR_API_KEY" \
  -d '{
    "url": "https://example.com/document.pdf",
    "store_name": "my-store"
  }'
```

#### Response `200 OK`

```json
{
  "success": true,
  "store_name": "fileSearchStores/my-store-31ihhh9rmyef",
  "file_name": "document.pdf",
  "message": "File uploaded and indexed successfully",
  "token_count": 4200,
  "estimated_cost_usd": 0.0021,
  "creditUtilisation": 2
}
```

#### Response 400 Bad Request&#xD;

```json
{
  "success": false,
  "message": "File type '.csv' is not supported. Allowed types: .pdf, .txt, .doc, .docx",
  "error": "File type '.csv' is not supported. Allowed types: .pdf, .txt, .doc, .docx"
}
```

#### Using the store in Chat Completions

Pass the `store_name` from the upload response into `file_search_store_names`:

**Request body:**

| Parameter                 | Type      | Required | Description                      |
| ------------------------- | --------- | -------- | -------------------------------- |
| `messages`                | array     | ✅        | Chat messages                    |
| `model`                   | string    | —        | Model ID                         |
| `file_search_store_names` | string\[] | —        | Store names from `uploadRagFile` |
| `thread_id`               | string    | —        | Conversation thread ID           |
| `user_id`                 | string    | —        | User identifier                  |
| `webhookUrl`              | string    | —        | Webhook for async results        |
| `metadata`                | object    | —        | Extra key/value metadata         |

```bash
curl -N https://qolaba-server-b2b.up.railway.app/api/v1/chat/completions
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemini-2.5-flash",
    "messages": [
      { "role": "user", "content": "Summarize the uploaded document" }
    ],
    "file_search_store_names": ["fileSearchStores/my-store-31ihhh9rmyef"],
    "stream": false
  }'
```

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/chat-completions-copy>" %}


# Chat

Bot API allows you to try out different Large Language Models (LLMs) from providers like Mistral, OpenAI, and Claude.

For more details on the supported LLM models and their capabilities, please refer to the [Broken mention](broken://pages/zF6PdCmIL9GuDtdhKuvb).

## Chat with LLM

<mark style="color:green;">`POST`</mark> `/chat`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

<table><thead><tr><th width="166">Name</th><th width="162">Type</th><th>Description</th></tr></thead><tbody><tr><td>llm</td><td>string</td><td><p>-> The <code>llm</code> parameter specifies the type of Large Language Model (LLM) to be used. The supported values are:</p><ul><li><code>ClaudeAI</code></li><li><code>OpenAI</code></li><li><code>GeminiAI</code></li><li><code>OpenRouterAI</code></li></ul><p></p></td></tr><tr><td>llm_model</td><td>string</td><td><p></p><p>-> The <code>llm_model</code> parameter specifies the name of the Large Language Model (LLM) to be used. The supported values for this parameter depend on the <code>llm</code> you have selected.</p><p>For example, if the <code>llm</code> is <code>OpenAI</code>, the supported <code>llm_model</code> values include:</p><ul><li><code>"gpt-4.1-mini-2025-04-14"</code></li><li><code>"gpt-4.1-2025-04-14"</code></li></ul><p>Similarly, if the <code>llm</code> is <code>ClaudeAI</code>, the supported <code>llm_model</code> values include:</p><ul><li><code>"claude-3-7-sonnet-latest"</code></li></ul><p>Ensure that the <code>llm_model</code> value you provide is compatible with the selected <code>model_type</code>.</p><p></p><p>-> Supported LLMs and their Models are : </p><p></p><p>GeminiAI</p><ol><li>gemini-2.5-pro </li><li>gemini-2.5-flash</li></ol><p></p><p>ClaudeAI</p><ol><li>claude-sonnet-4-20250514</li><li>claude-opus-4-1-20250805</li></ol><p></p><p>OpenAI</p><ol><li>gpt-4.1-mini-2025-04-14</li><li>gpt-4.1-nano-2025-04-14</li><li>gpt-4o-mini </li><li>gpt-4.1-2025-04-14</li><li>o1</li></ol><p>(Note: Below models for OpenAI LLM must be configured with <code>temperature=1</code> ,while runninng this model)</p><ol><li>o3-mini</li><li>o3</li><li>o4-mini-2025-04-16  </li><li>gpt-5-2025-08-07</li><li>gpt-5-mini-2025-08-07</li><li>gpt-5-nano-2025-08-07</li></ol><p></p><p>OpenRouterAI</p><ol><li>x-ai/grok-4</li><li>x-ai/grok-3-mini-beta</li><li>x-ai/grok-3-beta</li><li>perplexity/sonar-pro</li><li>perplexity/sonar-reasoning-pro</li><li>perplexity/sonar-reasoning</li><li>perplexity/sonar-deep-research</li><li>deepseek/deepseek-chat</li><li>deepseek/deepseek-r1</li></ol></td></tr><tr><td>history</td><td>List of Dictionary (Python) or map (C++)</td><td><p>-> The <code>history</code> parameter allows you to provide the previous chat history and the last user message. The history should follow a specific pattern:</p><ol><li>The history should start with a user message.</li><li>After each user message, there should be a response from the assistant.</li><li>The last message in the history should be a user message or query.</li></ol><p>-> The history should be provided as a list of dictionaries, where each dictionary represents a message. The dictionary should have the following structure:</p><pre><code> {
    "role": "user",
    "content": {
      "text": "please give 50 word story in English",
      "image_data": [
         {
           "url": "string",
           "details": "low"
         }
       ]
     }
  }

</code></pre><p>The <code>image\_data</code> parameter is a list that can contain up to 5 image URLs. The <code>details</code> parameter is optional and can be used to specify the level of detail required for the image analysis (only for OpenAI models).</p><p>Here's an example of the <code>history</code> parameter:</p><pre><code>"history": \[
{
"role": "user",
"content": {
"text": "hey",
"image\_data": \[]
}
},
{
"role": "assistant",
"content": {
"text": "Hello! How can I assist you today?",
"image\_data": \[]
}
},
{
"role": "user",
"content": {
"text": "Please analyze this image",
"image\_data": \[
{
"url": "<https://res.cloudinary.com/qolaba/image/upload/v1695690455/kxug1tmiolt1dtsvv5br.jpg>",
"details": "low"
}
]
}
}
]

</code></pre></td></tr><tr><td>temperature</td><td>float</td><td><p>-> The <code>temperature</code> parameter accepts a float value between 0 and 1. This parameter helps control the level of determinism in the output from the Large Language Model (LLM).</p><p>-> Higher temperature values (closer to 1) will result in more diverse and creative output, while lower values (closer to 0) will lead to more deterministic and conservative responses.</p></td></tr><tr><td>image\_analyze</td><td>bool</td><td>-> If you are passing image URLs and want the model to analyze the images, set the <code>image\_analyze</code> parameter to <code>true</code>.</td></tr><tr><td>enable\_tool</td><td>bool</td><td><p>-> To use the tools supported by the Chat API, enable the <code>enable\_tool</code> parameter. The Chat API currently supports two tools:</p><ol><li>Vector Search</li><li>Internet Search</li></ol><p>-> After enabling the <code>enable\_tool</code> parameter, you can provide the details of the tool you want to use in the <code>tools</code> parameter.<br></p></td></tr><tr><td>system\_msg</td><td>string</td><td>-> The <code>system\_msg</code> parameter allows you to set a system message for the Large Language Model (LLM). This message can be used to provide context or instructions to the model, which can influence the tone and behavior of the generated responses.</td></tr><tr><td>tools</td><td>dictionary or map</td><td><p></p><p>The <code>tools</code> parameter allows you to configure the capabilities available to the assistant during a conversation. It defines which tools are enabled, how much past context to consider, and specifies settings for embeddings, PDFs, and image generation. This is useful for tailoring the assistant’s functionality to your use case.</p><p><strong>🔹 Structure</strong></p><pre class="language-json"><code class="lang-json">jsonCopyEdit"tools": {
"tool\_list": {
"image\_generation": false,
"image\_generation1": false,
"image\_editing": false,
"search\_doc": false,
"internet\_search": false,
"python\_code\_execution\_tool": false,
"csv\_analysis": false
},
"number\_of\_context": 3,
"pdf\_references": \[],
"embedding\_model": \[
"text-embedding-3-large"
],
"image\_generation\_parameters": {}
} </code></pre><p><strong>🔹 Field Descriptions</strong></p><ul><li><p><strong>tool\_list</strong>: An object with boolean flags to enable or disable specific tools. Setting a value to <code>true</code> activates that feature.</p><ul><li><code>image\_generation</code>: Enables image creation from text prompts.</li><li><code>image\_generation1</code>: Optional alternate switch for image generation (if supported).</li><li><code>image\_editing</code>: Enables uploading and editing images.</li><li><code>search\_doc</code>: Allows document-based searching.</li><li><code>internet\_search</code>: Enables real-time web searches.</li><li><code>python\_code\_execution\_tool</code>: Allows execution of Python code within a safe sandbox.</li><li><code>csv\_analysis</code>: Enables processing and analysis of uploaded CSV files.</li></ul></li><li><strong>number\_of\_context</strong>: An integer defining how many previous user-assistant interactions to retain in memory. Higher values help preserve conversation flow.</li><li><strong>pdf\_references</strong>: A list of PDFs that the assistant can reference for contextual information should be stored in vector store database. Each entry includes an unique id which comes from the output of vector store API. If you're adding more than one document using <code>pdf\_references</code>, make sure you add the same number of entries in the <code>embedding\_model</code> array.</li><li><strong>embedding\_model</strong>: we support  <code>"text-embedding-3-large"</code> as  the embedding model.</li><li><strong>image\_generation\_parameters</strong>: An object to define settings like size or style for image generation. Leave it empty for default behavior.</li></ul></td></tr><tr><td>token</td><td>string</td><td>Authentication token needed for the request.</td></tr><tr><td>orgID</td><td>string </td><td>Identifies the organization associated with the request.</td></tr><tr><td>function\_call\_list</td><td>string</td><td>Lists the functions called during the process.</td></tr><tr><td>systemId </td><td>string</td><td>Specifies the system identifier linked to the request.</td></tr><tr><td>last\_user\_query</td><td>string</td><td>Provides the last user query made in the chat.</td></tr></tbody></table>

Example of Input body parameters :&#x20;

```
data '{
  "llm": "GeminiAI",
  "llm_model": "gemini-2.5-flash",
  "history": [
    {
      "role": "user",
      "content": {
        "text": "please give 50 word story in English",
        "image_data": [
          {
            "url": "string",
            "details": "low"
          }
        ]
      }
    }
  ],
  "temperature": 0.7,
  "image_analyze": false,
  "enable_tool": false,
  "system_msg": "You are a helpful assistant.",
  "tools": {
    "tool_list": {
      "image_generation": false,
      "image_generation1": false,
      "image_editing": false,
      "search_doc": false,
      "internet_search": false,
      "python_code_execution_tool": false,
      "csv_analysis": false
    },
    "number_of_context": 3,
      "pdf_references": [
      "output vectorstore1",
      "output vectorstore2"
    ],
    "embedding_model": [
      "text-embedding-3-large",
      "text-embedding-3-large"
    ],
    "image_generation_parameters": {}
  },
  "token": "123",
  "orgID": "string",
  "function_call_list": [],
  "systemId": "string",
  "last_user_query": "string"
}'
```

After passing the necessary parameters and executing the Chat API, you will receive a stream response. A successful response will look similar to the following:

**Response**

{% tabs %}
{% tab title="200" %}

<pre class="language-json"><code class="lang-json"><strong>{
</strong>  "output": null, 
  "promptTokens": null, 
  "completionTokens": null
}
</code></pre>

{% endtab %}

{% tab title="500" %}

```json
{
  "output": null, 
  "error": string, 
  "error_data": string
}
```

{% endtab %}
{% endtabs %}

The Chat API response is a streaming response, which means you will receive the output in chunks as the model generates the response. The response will continue to stream until the generation is complete.

The response will contain the following elements:

* `output`: This object contains the generated text output from the model.

The text output from the LLM can be obtained from the `output` parameter. When the response is complete, the final chunk will contain a `null` value in the `output` parameter, indicating the end of the stream.

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/chat-api>" %}


# Store File in Vector Database

## Parse File

Use this endpoint to add a PDF, CSV, TXT or DOC/DOCX document to the Qolaba AI database. The API will parse the document and store it in the vector database, returning a unique ID that can be used to retrieve information from the document using the Large Language Model (LLM).

Please note the following guidelines when indexing Document:

* The PDF, DOC/DOCX file should not exceed 200 pages.
* The CSV file should not contain more than 30 columns and 500 rows.
* When uploading a CSV file to the API, the first row must contain the column names. This helps the Large Language Model (LLM) better understand the values in each row of the CSV file.&#x20;
* Ensure that the document does not contain any sensitive or confidential information.

The unique ID returned after indexing the document can be used in subsequent requests to the Chatbot API to retrieve relevant information from the document, after ensuring enable\_tool and search\_doc to be true.

<br>

<mark style="color:green;">`POST`</mark> `/pdfVectorStore`

**Headers**

| Name          | Value              |
| ------------- | ------------------ |
| Content-Type  | `application/json` |
| Authorization | `Bearer <token>`   |

**Body**

| Name  | Type   | Description                                                          |
| ----- | ------ | -------------------------------------------------------------------- |
| `url` | string | The `url` parameter specifies the URL of the document to be indexed. |

**Response**

{% tabs %}
{% tab title="200" %}

```json
{
  "output": null, 
  "error": null, 
  "error_data": null
}
```

{% endtab %}

{% tab title="500" %}

```json
{
  "output": null, 
  "error": string, 
  "error_data": string
}
```

{% endtab %}
{% endtabs %}

Upon successfully indexing a document, the API will return a response with the unique identifier of the indexed document in `output` parameter.

You can use the unique identifier in subsequent requests to the Chat API to retrieve information from the indexed PDF.

## Run the API

To test this API, please use the following link:

{% embed url="<https://app.theneo.io/api-runner/qolaba/ml-apis/api-reference/store-file-in-vector-database>" %}


# Pricing&#x20;

## TEXT TO IMAGE

The credit utilization pricing for various models within the text-to-image feature is outlined below :-

<table><thead><tr><th width="312">App ID</th><th width="174">Model Name</th><th>Credit Utilisation</th></tr></thead><tbody><tr><td>ap-n2p3fg3gsvbgnYeEEdef</td><td>Image Gen 4</td><td>4</td></tr><tr><td>ap-nOpQr7stuvwxYzABcdef</td><td>SD 3.5 medium</td><td>9</td></tr><tr><td>ap-mNopQ8rstuvwXYZabcde</td><td>SD 3.5</td><td>23</td></tr><tr><td>ap-rStUv6xyzabcdPQRSefg</td><td>SD 3.5 Turbo</td><td>8</td></tr><tr><td>ap-jKlMn5opqzabcXyZtUVw</td><td>Recraft V3</td><td>22</td></tr><tr><td>ap-jXyZa9bcdefghijklmnopq</td><td>Flux Schnell</td><td>3</td></tr><tr><td>ap-fGhKl3mfkdlpqrtsUVWcba</td><td>Flux Dev</td><td>8</td></tr><tr><td>ap-fGhKl3mfkdlpqrsTuvwxYz</td><td>Flux Pro</td><td>14</td></tr><tr><td>ap-sdSyd0idsndjnsnsndjsds</td><td>Dalle 3</td><td>11</td></tr><tr><td>ap-x7q8hj9kltmNoPqRzabc</td><td>Leonardo</td><td>11</td></tr><tr><td>ap-hJkLm4nqzxybwvUTSRdca</td><td>Ideogram</td><td>22</td></tr><tr><td>ap-zuzhawbgipcrnxdtefhjbnvhc</td><td>GPT Image</td><td>18</td></tr></tbody></table>

### IMAGE TO IMAGE

The credit utilization pricing for various models within the image-to-image feature is outlined below :-

<table><thead><tr><th width="312">App ID</th><th width="174">Model Name</th><th>Credit utilisation  </th></tr></thead><tbody><tr><td>ap-fGhKl3mfkdlpqrshuwwabc</td><td>Flux Dev</td><td>8</td></tr><tr><td>ap-hTjXy7qplkzvnmwqrsdabc</td><td>Ideogram Remix</td><td>8</td></tr><tr><td>ap-gFhLq9zvbnmopxkytswabc</td><td>Flux Redux</td><td>26</td></tr><tr><td>ap-mNqZx4rtyjvbcfghwklpde</td><td>Flux Canny</td><td>26</td></tr><tr><td>ap-gWfZx8rjvtyqopmnbcdehij</td><td>Flux Depth</td><td>26</td></tr><tr><td>ap-hJkLm8nqwertyuiopasdfg</td><td>Flux PulID</td><td>17</td></tr><tr><td>ap-zuzhawbgipcrnxdtefhjbnvhc</td><td>GPT Image</td><td>18</td></tr><tr><td>ap-tK7gR4jP9sWbC1mNfE8uVzH2</td><td>Flux Kontext Max</td><td>29</td></tr></tbody></table>

#### INPAINTING

The credit utilization pricing for INPAINTING within the image-to-image feature is 10.

#### REPLACE BACKGROUND

The credit utilization pricing for REPLACE BACKGROUND within the image-to-image feature is 10.

#### TEXT TO SPEECH

The credit utilization pricing for various voice models within the TEXT TO SPEECH feature is 2.


# Overview

An overview of how Organizations and Workspaces work in Qolaba — credit management, collaboration, access control, and usage visibility.

### **Why Organizations & Workspaces Exist**

Qolaba is designed to work seamlessly for both individuals and teams. As your usage grows—across projects, clients, or departments—you need a structured way to manage credits, organize work, and collaborate efficiently.

**Organizations and Workspaces provide that structure.**

They help you move from isolated usage to a more **organized, collaborative, and scalable setup**, without losing visibility or control.

***

### Key Challenges Addressed

<table><thead><tr><th width="188.02032470703125">Challenge</th><th>Without Structure</th><th>With Organizations &#x26; Workspaces</th></tr></thead><tbody><tr><td>Credit management</td><td>Scattered usage across users</td><td>Centralized credit pool or Separate Workspace Budgets</td></tr><tr><td>Collaboration</td><td>Disconnected workflows</td><td>Shared environment</td></tr><tr><td>Access control</td><td>No clear permissions</td><td>Role-based access</td></tr><tr><td>Organization</td><td>Mixed outputs and history</td><td>Structured workspaces</td></tr><tr><td>Visibility</td><td>Difficult to track usage</td><td>Clear usage tracking</td></tr></tbody></table>

***

### Core Capabilities

#### **1. Centralized Credit Management**

All usage in Qolaba is credit-based. Instead of managing credits individually, organizations introduce a **shared system**:

* A single subscription and billing source
* Allocate budgets to each workspace or consume from shared pool
* Centralized top-ups and plan management
* Visibility into overall usage

This makes it easier to control costs and avoid fragmented credit usage across multiple users.

***

#### **2. Structured Collaboration**

Organizations and Workspaces allow teams to collaborate in a structured and organized way, while keeping work relevant to each role.

* Different teams or functions can work within their own **dedicated workspaces**
* Each workspace can focus on specific tasks or workflows
* Usage and outputs remain organized and easy to track

{% hint style="info" %}
**Example: Team Workflow**

A team organizes work across multiple workspaces:

* In the **Research workspace**:\
  Members use **Chat (LLMs)** to generate insights, content ideas, and strategy inputs
* In the **Design workspace**:\
  Members use **Image and Video tools** to create visuals and assets
* In the **Content workspace**:\
  Members refine messaging, scripts, and final outputs

Each team works within its own context, while all activity stays connected under the same organization.
{% endhint %}

***

#### **3. Clear Separation of Work**

As more work is created, keeping everything in one place becomes difficult to manage.

Workspaces help you structure work based on your needs, such as:

* Projects or clients
* Teams or departments
* Use cases or workflows

Each workspace acts as an independent environment, helping you keep work cleanly separated without affecting other teams.

***

#### **4. Usage Visibility**

Organizations provide a clear view of how credits are being used across your team.

You can:

* Monitor usage across different workspaces
* Track activity at a member level
* Analyze usage across tools (Chat, Image, Video, Audio)

This level of visibility helps teams stay **in control of usage and spending**.

***

### Summary

Organizations and Workspaces form the foundation for using Qolaba effectively at scale.

They allow you to:

* Manage credits from a single place
* Collaborate without overlap
* Keep work structured and easy to navigate
* Maintain visibility into usage

Whether you're working solo or as part of a team, this structure ensures your workflows remain **organized, flexible, and scalable**.

***

### Suggested Next Steps

To learn more, explore:

* [Creating an Organization](/organization-and-workspaces/organizations/creating-an-organization)
* [Creating and Managing Workspaces](/organization-and-workspaces/workspaces)
* [Inviting and Managing Members](/getting-started/introduction)
* [Credits and Usage Tracking](/organization-and-workspaces/tracking-credits-and-ai-usage)


# Organizations

Understand how Organizations help manage teams, workspaces, billing, and credits in one centralized place.

An **Organization** is the central layer in Qolaba where your team, credits, and workspaces are managed together.

It acts as the **single source of truth** for:

* Billing and credit usage
* Team members and access
* Workspace structure

If you're working with multiple people, projects, or clients, creating an organization allows you to keep everything **centralized, organized, and controlled**.

***

### When to Use an Organization

You should use an organization when:

* You’re working as a **team** and need shared access
* You want to manage **credits centrally** instead of individually
* You need to organize work across **multiple projects or clients**
* You want **visibility into usage** across members and workspaces

***

### What You Can Do with an Organization

Within an organization, you can:

* Create and manage multiple workspaces
* Invite and manage team members
* Control access and permissions
* Monitor credit usage and activity
* Manage billing and subscriptions

This allows you to operate in a **structured and scalable way**, without losing control as your usage grows.

***

### How Organizations Fit Into Your Workflow

Once an organization is created:

1. You set up your **workspaces** based on projects or teams
2. Invite **members** and assign them access
3. Start using AI tools within each workspace
4. Track usage and manage credits centrally

This ensures that all activity remains **organized and easy to manage**.


# Creating an Organization

Learn how to create an Organization in Qolaba, choose the right credit plan, and complete the initial setup for teams, billing, and workspaces.

Creating an organization is the first step to setting up Qolaba for structured usage. It allows you to manage credits centrally, collaborate with others, and organize your work in a scalable way.

An organization acts as the main container for:

* Subscription and billing
* Credit allocation
* Workspaces
* Team members

***

### **1. Access Organization Setup**

* **Step 1**\
  Go to your **Qolaba Dashboard**
* **Step 2**\
  Click your **profile icon** (top-right corner)
* **Step 3**\
  Select **Create Organization**

***

### **2. Organization Setup Flow**

The setup process is guided and takes just a few steps:

* Name your organization
* Select a credit plan
* Add basic details
* Complete checkout

#### **2.1 Enter Organization Name**

Provide a name that represents how you plan to use Qolaba.

{% tabs %}
{% tab title="Company / Brand" %}

* Growth Studio
* FoodCo
* Marketech
  {% endtab %}

{% tab title="Team / Department" %}

* Marketing
* Product
* Research
  {% endtab %}

{% tab title="Project-Based" %}

* Project Alpha
* Campaign Launch
* Internal Experiments
  {% endtab %}
  {% endtabs %}

***

#### **2.2 Select a Credit Plan**

Choose a plan based on how you expect to use the platform.

Each plan clearly shows:

* Number of credits included
* Pricing

**How to Choose the Right Plan**

Consider the following:

* **Team size**\
  More users typically require more credits
* **Type of work**
  * Text-based tasks → lower usage
  * Image and video generation → higher usage
* **Frequency of usage**\
  Occasional vs. daily workflows
* **Nature of work**
  * Lightweight exploration
  * High-volume production

{% hint style="info" %}
Credits are shared across the organization either through shared pool or through allocated workspace budgets.&#x20;

[Learn More About Workspace Credit Allocation *→*](/organization-and-workspaces/workspaces/workspace-credit-allocation)
{% endhint %}

***

#### **2.3 Provide Organization Details**

Add a few details to complete setup.

* **Company size range** (e.g., 2–50, 50–250, etc.)
* **Your role** (e.g., Marketing, Product, Research, Operations)

{% hint style="info" %}
&#x20;This helps tailor onboarding and recommendations based on how you plan to use Qolaba.
{% endhint %}

***

### **3. Complete Checkout**

* **Step 1**\
  Click **Create Organization & Continue**
* **Step 2**\
  Proceed to the **payment page**
* **Step 3**\
  Complete your subscription

***

### **After Setup**

Once your organization is created:

* Your **subscription becomes active**
* Credits are **allocated immediately**
* You are redirected to the **Organization Dashboard**
* You are assigned the **Owner role**

***

### **What You Can Do Next**

From your organization dashboard, you can:

* [**Create** **workspaces**](/organization-and-workspaces/workspaces/creating-a-workspace) to organize your work
* [**Assign Budgets**](/organization-and-workspaces/workspaces/workspace-credit-allocation) to each workspace
* [**Invite team members**](/organization-and-workspaces/members-and-role-based-access/inviting-members) <sub>**(only in team plan)**</sub>
* [**Monitor Credit usage**](/organization-and-workspaces/tracking-credits-and-ai-usage)


# Organization Dashboard

{% content-ref url="/pages/n8mXF6ZOO3zqP66EcqAu" %}
[Overview Tab](/organization-and-workspaces/organizations/organization-dashboard/overview-tab)
{% endcontent-ref %}


# Overview Tab

The Overview tab provides a complete summary of your organization’s financial status, credit availability, and usage activity. It combines:

* Organization summary
* Credit controls
* Billing records
* Credit consumption tracking

This is where Owners and Admins monitor and manage AI resources.

***

## 1. Organization Summary

The Organization Summary provides a high-level snapshot of your organization’s AI capacity and team structure. It typically includes:

* Organization name
* Monthly credit allocation
* Remaining credits
* Total members

Each element serves a specific purpose.

***

### Organization Name

This displays the official name of your organization. It:

* Identifies the team environment you are currently operating in
* Helps users confirm they are inside the correct organization
* For organizations managing multiple teams or clients, this helps prevent confusion when switching between environments.

***

### Monthly Credit Allocation

This shows the total number of credits assigned to your organization for the current billing cycle.This value is determined by:

* The subscription plan selected during setup
* Any upgrades made afterward

Credits are refreshed based on your billing cycle.This allocation represents the maximum AI usage capacity for the month.It defines your operational limit for:

* Chatbot usage
* Image generation
* Video generation
* Speech generation
* Knowledge base uploads

This helps teams plan resource usage strategically.

***

### Remaining Credits

This shows how many credits are still available for the current billing period.Credits decrease when:

* Any member uses AI tools
* Content is generated
* Files are processed

Remaining credits update in real time.This metric is critical for:

* Budget monitoring
* Preventing unexpected exhaustion
* Deciding whether to top up credits
* Planning large campaigns

For example:If you are launching a marketing campaign with heavy image and video generation, monitoring remaining credits prevents workflow disruption.

***

### Total Members

This shows the number of members currently part of the organization.It provides visibility into:

* Team size
* Active collaboration scale
* Distribution of credit consumption

This does not display detailed member information here — only the count.It helps leadership understand team footprint within the AI environment.

***

## 2. Credit Management

The Credit Management section allows Owners and Admins to control subscription capacity without leaving the dashboard.It includes two main controls:

* Buy More Credits
* Upgrade Plan

***

### Buy More Credits (Top-Up)

The “Buy More Credits” option allows organizations to instantly add credits to their existing subscription.This is useful when:

* Credits are running low mid-cycle
* A large campaign requires additional AI usage
* A new project begins unexpectedly
* A temporary spike in usage occurs

Top-ups:

* Increase available credits immediately
* Do not change your base subscription tier
* Act as supplemental credit additions

This provides operational flexibility without requiring a full plan upgrade.It prevents downtime and ensures uninterrupted AI workflows.

***

### Upgrade Plan (Change Subscription Tier)

Upgrading your plan increases your monthly credit allocation.This is ideal when:

* Your team has grown
* AI usage has permanently increased
* You consistently exhaust monthly credits
* You need more predictable scaling

Upgrading adjusts:

* Monthly credit allowance
* Subscription cost
* Long-term usage capacity

Unlike top-ups (which are temporary), upgrading changes your recurring subscription level.This is suited for sustainable growth rather than temporary spikes.

***

## 3. Billing Section

The Billing section provides financial transparency and subscription tracking.This is essential for:

* Accounting teams
* Budget reporting
* Finance audits
* Enterprise compliance

***

### Billing History

Billing History records every subscription and payment event.Each entry includes:

#### Date & Time

When the transaction occurred.Useful for:

* Monthly reconciliation
* Tracking subscription cycles

***

#### Plan Name

The name of the subscription plan purchased or renewed.Helps track:

* Plan changes over time
* Upgrade or downgrade history

***

#### Credit Quantity

The number of credits associated with the transaction.This confirms:

* Allocation received
* Top-up quantity purchased

***

#### Payment Amount

The total amount charged.Used for:

* Financial reporting
* Cost tracking
* Expense management

***

#### Subscription Status

Indicates whether the subscription is:

* Active
* Cancelled
* Expired
* Pending

This ensures leadership knows the operational state of the organization.

***

## 4. Credit Utilization Section

While Billing shows money flow, Credit Utilization shows operational AI activity.This section provides insight into how credits are being used across the platform.

***

### Platform-Wide Usage

This table records AI consumption events across all tools.Each entry includes:

***

#### Date & Time

When the AI action occurred.Helps track:

* Usage spikes
* Campaign periods
* High-activity windows

***

#### Tool Type

Specifies which tool consumed credits, such as:

* Chatbot
* Text-to-Image
* Image-to-Image
* Image Editing
* Video Generation
* Speech Generation
* Text-to-Music
* Knowledge Base processing

This helps identify which tools drive the highest usage.For example:If Video Generation consumes most credits, you may adjust strategy accordingly.

***

#### Quantity of Credits Used

Shows how many credits were consumed for that specific action.This enables:

* Cost accountability
* Usage analysis
* Optimization decisions
* Budget forecasting

***


# Profile Tab

## Profile Tab

The Profile tab is the administrative settings area of your organization.It controls:

* Organization identity
* Structural management
* Subscription lifecycle

Access to this tab depends on role:

* **Owner** → Full access
* **Admin** → Can edit most settings
* **Member** → No access

***

## 1. Organization Information

The Organization Information section defines the identity of your organization within Qolaba.It includes:

* Organization logo
* Organization name

This section allows Owners and Admins to customize and maintain the organization’s profile.

***

### Upload Organization Logo

You can upload a logo to visually represent your organization.

#### Purpose of Organization Logo

* Displays branding across the organization dashboard
* Improves team identification
* Helps members distinguish between multiple organizations
* Strengthens professional presentation

For agencies managing multiple organizations or consultants working with multiple clients, this helps avoid confusion when switching between accounts.The logo is purely organizational branding — it does not affect credit usage or functionality.

***

### Edit Organization Name (Owner & Admin Only)

Owners and Admins can modify the organization name at any time.

#### Why Editing May Be Needed

* Rebranding
* Company name change
* Department restructuring
* Project-to-company transition

The organization name:

* Appears in member invitations
* Appears in profile dropdown
* Appears in dashboard headers

#### Role Restriction

Members cannot edit the organization name.This ensures identity consistency and prevents accidental renaming.

***

## 2. Organization Management

This section controls the structural existence of the organization.It includes one critical action:Delete Organization

***

### Delete Organization

This action permanently removes the organization from Qolaba.

#### Access Restriction

Only the **Owner** can delete an organization.Admins cannot delete the organization.This ensures:

* Maximum structural security
* Protection from accidental deletion
* Clear authority boundary

***

#### What Happens When an Organization Is Deleted

Deleting an organization will:

* Remove all workspaces
* Remove all members
* Remove usage history
* Remove billing association
* Permanently erase associated data

This action **cannot be undone**.Once deleted:

* Credits are lost
* Workspace history is removed
* Member access is revoked
* Data cannot be recovered

***

#### When Should You Delete an Organization?

* Company closure
* Project termination
* Migration to a new organization
* Complete discontinuation of service

For temporary discontinuation, canceling subscription is safer than deleting.

***

## 3. Subscription Management

This section controls the lifecycle of your paid plan.It allows Owners and Admins to cancel active subscriptions.

***

### Cancel Subscription

Owners and Admins can cancel the organization’s subscription.Canceling subscription:

* Stops automatic renewal
* Does not immediately disable access

***

#### What Happens After Cancellation

If you cancel:

* Your subscription remains active until the end of the current billing cycle
* Credits remain usable during that period
* After expiration, premium access is discontinued

This ensures:

* No sudden workflow interruption
* Predictable offboarding
* Smooth transition planning

***

#### Important Distinction

Cancel Subscription ≠ Delete OrganizationCanceling:

* Stops future billing
* Keeps organization structure intact
* Ends premium features after billing period

Deleting:

* Permanently removes organization and all data

###


# Members Tab

## Members Tab

The Members tab is the administrative control center for managing users inside an organization.This tab is visible only to:

* Owner
* Admin

Members cannot access this tab.It allows leadership to:

* Monitor individual credit usage
* Assign and modify roles
* Manage workspace access
* Revoke permissions
* Maintain security and accountability

***

## 1. Member Overview Table

The Member Overview Table provides a structured summary of all invited users within the organization.Each row represents one member.Each column provides operational and governance insights.

***

### Email

Displays the registered email address of the member.This is:

* The email used to log into Qolaba
* The identifier for workspace access
* The email used for invitation

Why this matters:

* Ensures correct identity verification
* Prevents duplicate access
* Provides clarity on who is consuming credits

***

### Image & Speech Credits Used

This column shows how many credits the member has consumed across:

* Text-to-Image
* Image-to-Image
* Image Editing
* Video generation
* Speech generation
* Text-to-Music

This helps administrators understand:

* Who is generating creative assets
* Which team members are heavy media users
* Whether usage aligns with responsibilities

For example:If a copywriter is consuming high image credits, this may indicate cross-role usage or inefficiency.

***

### Chatbot Credits Used

This column shows credits consumed specifically for:

* Chat interactions
* Knowledge base usage
* AI reasoning tools

This separation is important because:

* Chatbot usage patterns differ from media generation
* Teams may allocate chatbot usage differently
* Some roles rely heavily on conversational AI

This enables better tool-level monitoring.

***

### Total Credits Used

This column shows the total credits consumed by the member across all AI tools.This provides:

* A complete usage snapshot
* Accountability per user
* Budget visibility
* Usage comparison between team members

This is especially useful for:

* Large teams
* Enterprise governance
* Client-billable environments

***

### Status (Active / Revoked)

This shows the current access status of the member.

#### Active

The member:

* Can log in
* Can access assigned workspace
* Can consume credits

#### Revoked

The member:

* No longer has access
* Cannot log into the organization
* Cannot consume credits

Revoking access does not delete usage history — it only removes permission.This is important for:

* Team transitions
* Employee offboarding
* Contractor termination
* Security enforcement

***

### Role

This shows the assigned role of the member:

* Owner
* Admin
* Member

This column clarifies:

* Who has administrative authority
* Who can manage credits
* Who can modify workspaces
* Who has restricted access

This ensures transparency in organizational hierarchy.

***

### Workspace

This column shows which workspace the member is assigned to.Because workspaces isolate projects or teams, this helps administrators understand:

* Which project the member belongs to
* Which client they are working on
* Where their AI activity is recorded

This prevents cross-project confusion.

***

## 2. Filtering Members

As organizations scale, the number of members increases.Filtering ensures administrative clarity.

***

### View All Workspaces

When “All” is selected:

* You see all members across all workspaces
* Full organizational visibility

This is useful for:

* Organization-wide audits
* Budget review
* High-level reporting

***

### Filter by Specific Workspace

You can filter members by selecting a specific workspace.When filtered:

* Only members assigned to that workspace appear
* Credit usage reflects that context

This is especially useful for:

* Client-level monitoring (agencies)
* Department-level tracking (SMBs)
* Project-level billing (consultants)

Filtering ensures clean separation without mixing unrelated teams.

***

## 3. Managing Members

The Members tab is not just for viewing — it allows administrative actions.

***

### Change Role

Admins and Owners can modify a member’s role.For example:

* Promote Member → Admin
* Demote Admin → Member

Why this is important:

* Team restructuring
* Temporary elevated permissions
* Project leadership changes
* Security adjustments

Role changes immediately affect dashboard access and permissions.

***

### Change Workspace

You can reassign a member to a different workspace.This is useful when:

* Team members shift projects
* Client ownership changes
* Departments reorganize
* A contractor finishes one project and moves to another

Workspace reassignment ensures:

* Correct history isolation
* Accurate credit tracking
* Clean project boundaries

***

### Revoke Access

Revoking access removes the member from the organization.This is used for:

* Employee exits
* Contract completion
* Security incidents
* Temporary suspension

When revoked:

* The member loses access immediately
* Credits are no longer consumable
* Usage history remains for reporting

&#x20;

This maintains operational integrity while preserving historical records.


# History Tab

The History tab allows users to view AI activity across the platform.It provides structured visibility into:

* What was generated
* When it was generated
* Which workspace it belongs to
* What type of content it is
* Whether it was private or public

Access level depends on role:

* **Owner & Admin** → Can view history across all workspaces
* **Member** → Can view history only within assigned workspace

***

## 1. Viewing Platform Activity

The History tab serves as a centralized activity feed.It allows you to:

* Track AI usage
* Revisit past outputs
* Monitor content generation
* Audit credit consumption patterns

Each entry in History represents an AI action, such as:

* Chat interaction
* Image generation
* Video creation
* Audio generation
* Uploaded file
* Edited content

This creates a full operational timeline of your organization’s AI usage.

***

## 1. 1 Enable “Me Mode”

“Me Mode” filters history to show only your own activity.When enabled:

* You see only content generated by your account
* Other team members’ activity is hidden

#### Why Me Mode Exists

For organizations with large teams, history can become extensive.Me Mode allows:

* Personal productivity tracking
* Clean workflow management
* Focused review of your own outputs
* Reduced noise from other team activity

For Members, this is especially useful when:

* They only need to revisit their own generated content
* They want to isolate their own work

For Admins, it helps compare personal usage against total team usage.

***

## 1.2 Filter by Workspace

The workspace filter allows you to isolate activity based on project or team.

***

### Primary Workspace (Default)

When an organization is created, the first workspace created is called the Primary Workspace.By default:

* History shows activity from the Primary Workspace

This ensures that even new organizations have a visible starting point.

***

### Switching Between Workspaces

If multiple workspaces exist, you can switch between them.When you select a workspace:

* History updates to show only activity from that workspace
* Content from other workspaces is hidden

#### Why This Matters

Workspaces isolate projects.For example:Agency:

* Workspace A → Client A
* Workspace B → Client B

Filtering by workspace ensures:

* Client content does not mix
* Clean project-level tracking
* Accurate billing attribution

Members can only view the workspace they are assigned to.Admins can switch between all workspaces.

***

## 1.3 Sort by Date

Sorting controls how activity is displayed chronologically.

***

### Latest First

Displays the most recent activity at the top.Best used for:

* Daily monitoring
* Reviewing recent work
* Tracking new outputs
* Ongoing campaign management

This is typically the default view.

***

### Oldest First

Displays the earliest activity first.Best used for:

* Reviewing long-term project timelines
* Auditing historical usage
* Tracking workflow progression
* Investigating past campaigns

Sorting ensures flexible time-based analysis.

***

## 2. Filter by Content Type

Beyond workspace and user filtering, you can refine history by content category.This allows precision search across generated assets.

***

### Uploaded Images

Displays images that were manually uploaded to the platform.Useful for:

* Locating base images
* Revisiting reference uploads
* Reusing previous assets

***

### Image Content

Displays AI-generated images.Includes:

* Text-to-Image
* Image-to-Image
* Image Editing outputs

Helps users quickly locate visual outputs.

***

### Audio Content

Displays:

* Speech generation
* Text-to-Music outputs
* Audio files created via AI

Useful for creators and media teams working with sound assets.

***

### Video Content

Displays:

* Generated videos
* Image-to-video outputs
* Text-to-video creations

This is especially useful for:

* Campaign review
* Video asset tracking
* Creative production oversight

***

### Private Content

Shows content generated under private session mode.Private content:

* Is visible only to the creator
* Does not appear on the community wall

This filter helps isolate confidential or client-sensitive work.

***

### Public Content

Shows content generated under public mode.Public content:

* Is visible on the community wall
* May be shareable

This filter helps separate:

* Marketing/public campaigns
* Showcase content
* Collaborative outputs

***

### Favorites

Displays content that has been marked as favorite.This is useful for:

* Shortlisting final creatives
* Saving best-performing assets
* Organizing priority outputs
* Quick access to important files

Favorites act as a lightweight bookmarking system.

{% hint style="info" %}

{% endhint %}


# Workspaces

Understand how Workspaces help organize projects, teams, and AI usage within an Organization.

Workspaces are structured environments inside an organization that allow teams to separate projects, clients, or departments while sharing the same subscription. They act as controlled containers for AI usage. Workspaces can access credits directly from the shared pool or through their allocated budgets.

### What is a Workspace?

A workspace is a **project-level container** within an organization designed to structure and manage AI operations.

It is used to:

* Organize AI work by project, team, or client
* Maintain separation between different use cases
* Track credit usage independently
* Preserve clean and isolated activity history

#### Common Use Cases

| User Type       | Workspace per... | Example                       | Benefit                       |
| --------------- | ---------------- | ----------------------------- | ----------------------------- |
| Agency          | Client           | Workspace: Nike Campaign      | Data isolation, clean billing |
| SMB             | Department       | Workspace: Marketing          | Usage tracking per team       |
| Consultant      | Engagement       | Workspace: Q3 Audit           | Client confidentiality        |
| Freelancer      | Revenue stream   | Workspace: Brand A            | Isolated outputs              |
| Dev / R\&D Team | Experiment       | Workspace: GPT vs Claude Test | No production interference    |

***

### Key Capabilities

#### 1. Project or Client-Level Structuring

A workspace can represent different operational units such as:

* Clients
* Departments
* Campaigns
* Product lines
* Research initiatives
* Consulting engagements

This prevents overlap between unrelated work and ensures structured execution.

#### 2. Independent Usage Tracking

Although credits belong to the organization, usage is tracked at the workspace level.

Each workspace provides visibility into:

* Image and speech credit consumption
* Chatbot credit usage
* Total credits consumed

This enables accurate:

* Client billing
* Department budgeting
* Resource forecasting
* Performance analysis

#### 3. Isolated AI History

Every workspace maintains its own activity history.

This ensures:

* No cross-workspace visibility of generated content
* Project-specific conversations remain contained
* Outputs and media are properly organized

This isolation supports confidentiality, privacy, and clean workflows.

***

### Primary Workspace

When an organization is created, **Qolaba automatically generates the first workspace**, known as the **Primary Workspace**.

The Primary Workspace:

* Is the first workspace created by default
* Serves as the initial working environment

This ensures the organization is immediately ready for use without additional setup.

***

#### Default Workspace Behavior

The Primary Workspace acts as the default workspace for all initial activities.

This means:

* All initial AI activity is recorded here
* History is stored by default in this workspace
* Members operate within it unless reassigned

If no additional workspaces are created, all activity remains within the Primary Workspace.

#### When Should You Create Additional Workspaces?

You should create new workspaces when:

* Managing multiple clients
* Operating across different departments
* Requiring separate credit tracking per project
* Needing isolated AI history
* Enforcing stricter project-level governance

Workspaces are essential for maintaining scalability, clarity, and operational control.

***

### Organization vs Workspace

<table><thead><tr><th width="136.84375">Aspect</th><th>Organization</th><th>Workspace</th></tr></thead><tbody><tr><td>Role</td><td>Top-level entity</td><td>Project-level container</td></tr><tr><td>Ownership</td><td>Owns subscription and credits</td><td>Uses allocated or shared credits</td></tr><tr><td>Management</td><td>Manages members</td><td>Assigns members per project</td></tr><tr><td>Function</td><td>Contains workspaces</td><td>Segments and organizes work</td></tr><tr><td>Data Scope</td><td>Global</td><td>Isolated per workspace</td></tr></tbody></table>


# Creating a Workspace

Learn how to create and configure Workspaces for projects, teams, or clients within your Organization.

### Steps to Create a Workspace

1. Go to the **Organization Dashboard**
2. Open the **Workspaces** tab
3. Click **Create Workspace**
4. Enter the required details
5. Click **Create Workspace**

Once created, the workspace becomes immediately active and ready for use.

{% hint style="info" %}
Only users with **Owner** or **Admin** roles can create workspaces.\
[Learn More about Role Based Acces&#x73;*→*](/organization-and-workspaces/members-and-role-based-access/member-roles-explained)
{% endhint %}

***

### Workspace Details

#### 1. Workspace Name

The workspace name identifies the project, team, or client within the organization.

Use clear and descriptive naming conventions such as:

* Client – Nike Campaign
* Marketing Team
* Product Launch 2024
* Q4 Ad Campaign

The name is used across member assignment, history filters, and credit tracking. A consistent naming approach ensures clarity as the number of workspaces increases.

{% hint style="info" %}
Prefer descriptive names over generic ones\
*Example: “Client – Adidas Q4 Campaign” is more effective than “Workspace 3”*
{% endhint %}

#### 2. Workspace Description

You can add a description to provide context about the workspace.

This may include:

* Project scope or objective
* Client or campaign details
* Timeline or phase
* Responsible team

**Example:**“This workspace is dedicated to AI-generated creatives and chatbot workflows for the Q3 product launch.”

***

### After Creation

Once the workspace is created, you can:

* [Invite members](/organization-and-workspaces/members-and-role-based-access/inviting-members) to the workspace
* [Allocate a budget](/organization-and-workspaces/workspaces/workspace-credit-allocation) to the workspace (if using workspace budget mode)
* [Track credit usage](/organization-and-workspaces/tracking-credits-and-ai-usage)
* Access workspace-specific history
* Update workspace details
* Delete the workspace if no longer needed

All activity within the workspace is automatically tracked and scoped to it.

***

### Best Practices

* Create separate workspaces for distinct clients, teams, or projects
* Maintain consistent naming conventions across the organization
* Avoid mixing unrelated work within a single workspace
* Do not create unnecessary workspaces without clear separation needs


# Workspace Credit Allocation

Understand how credit allocation works across Workspaces using shared pools or individual workspace budgets.

Workspace Credit Allocation defines how credits are distributed across workspaces. It allows organizations to either use a shared credit pool or assign fixed budgets to individual workspaces, enabling controlled and transparent usage.

{% embed url="<https://youtu.be/r6NShDw7Nls?si=wEBv0NmPeYNlex0D>" %}

***

### Credit Modes

You can switch between modes from the dropdown at the top of the **Workspaces** tab.

#### 1. Shared Pool Mode

All workspaces consume credits from a single shared pool.

* Credits reside in the **Primary Workspace**
* No per-workspace limits are enforced
* All usage deducts from the same central balance

**When to use:**\
Suitable for smaller teams or when centralized control is sufficient.

#### 2. Workspace Budget Mode

Each workspace operates within an assigned credit limit.

* Credits are distributed to individual workspaces
* Usage is restricted to the allocated budget
* Additional credits must be assigned when required

**When to use:**\
Recommended for managing multiple teams, clients, or projects with defined budgets.

***

### Primary Workspace

The **Primary Workspace** acts as the default credit source.

* All purchased or added credits are stored here
* Its balance updates dynamically based on allocations

**Example:**\
If total credits = 10,000

* 2,000 allocated to Workspace A
* 3,000 allocated to Workspace B

Primary Workspace balance = **5,000**

***

### Allocating Credits

Credits can be transferred between workspaces in **Workspace Budget Mode**.

#### Steps

1. Go to **Dashboard → Workspaces tab**
2. Open the **⋯ (three-dot menu)** for a workspace
3. Select **Manage Credits**
4. Choose a source workspace
5. Enter the credit amount
6. Review updated balances
7. Click **Transfer**

***

### Credit Behavior

* Each workspace consumes credits from its own balance (in budget mode)
* The Primary Workspace reflects the remaining organizational balance
* If a workspace runs out of credits, usage is blocked until more credits are assigned

***

### Transfer Rules

* Credits can be moved between any workspaces with available balance
* Transfers cannot exceed available credits in the source workspace
* Changes apply immediately

***

### Deletion Impact

* Remaining credits in a deleted workspace return to the **Primary Workspace**
* The Primary Workspace cannot be deleted

***

### Why This Matters

* Enables controlled distribution of credits
* Prevents uncontrolled usage
* Provides visibility at the workspace level
* Ensures accountability across teams


# Managing Workspaces

Learn how to monitor workspace usage, manage settings, and perform administrative actions from the Organization Dashboard.

Workspaces can be monitored and managed from the **Workspaces** tab in the Organization Dashboard. This section focuses on viewing workspace activity and performing administrative actions.

***

### Workspace View

The Workspaces tab provides a structured overview of all workspaces.

Each row includes:

* **Workspace Name** — Project, team, or client identifier
* **Image & Speech Credits Used** — Media-related usage
* **Chatbot Credits Used** — Text and conversation usage
* **Total Credits Used** — Combined usage across all tools
* **Usage Display**
  * Shows consumption against allocation (in budget mode)
  * Shows total usage from shared pool (in shared mode)
* **Created By** — Workspace creator

This view helps compare usage and monitor activity across workspaces.

***

### Understanding Usage

Workspace metrics help interpret how resources are being used:

* Higher **Image & Speech usage** indicates creative or production-heavy work
* Higher **Chatbot usage** indicates research or conversational workflows
* **Total usage** reflects the overall resource consumption of the workspace

***

### Workspace Actions

Each workspace provides the following actions:

* **Edit** — Update workspace name or description
* **Delete** — Remove the workspace

If **Workspace Budget Mode** is enabled:

* **Manage Credits** — Opens credit allocation controls (covered in [Workspace Credit Allocation](/organization-and-workspaces/workspaces/workspace-credit-allocation))

{% hint style="info" %}
Only users with **Owner** or **Admin** roles can create workspaces.\
[Learn More about Role Based Acces&#x73;*→*](/organization-and-workspaces/members-and-role-based-access/member-roles-explained)
{% endhint %}

***

### Updating a Workspace

Workspace details can be modified to reflect:

* Project updates
* Naming changes
* Organizational restructuring

These changes do not affect usage data or member assignments.

***

### Deleting a Workspace

Deleting a workspace:

* Removes the workspace and its activity scope
* Revokes workspace-specific access
* Stops further usage

Credits already consumed remain part of organizational usage.

{% hint style="warning" %}
Deleted workspaces cannot be recovered. Archive or export relevant content before deletion.
{% endhint %}


# Members and Role-Based Access

Understand how member roles and workspace-based access help manage collaboration, permissions, and security within your Organization.

## Members & Role-Based Access

Qolaba uses a structured, role-based access system to help organizations collaborate securely, manage credits transparently, and keep workspaces isolated. This section covers everything you need to know about managing who has access to your organization — and what they can do once they're in.

***

### What's in This Section

<table><thead><tr><th width="240.57421875">Page</th><th>Description</th></tr></thead><tbody><tr><td><a href="/pages/u0FWrK29aw1Xkwm5P5Tv">Inviting Members</a></td><td>How to add team members, assign roles, and select workspaces</td></tr><tr><td><a href="/pages/ty7QrDf1saw287GvEayZ">Member Activation Flow</a></td><td>What a new member does after receiving an invitation</td></tr><tr><td><a href="/pages/AGJQlT2d3CiVRO7RFJys">Roles Explained</a></td><td>Full breakdown of Owner, Admin, and Member permissions</td></tr></tbody></table>

***

### How It Works — At a Glance

Every Qolaba organization has three role tiers:

* **Owner** — Full control, including the ability to delete the organization
* **Admin** — Manages members, workspaces, and credits; cannot delete the organization
* **Member** — Uses AI tools within their assigned workspace; no administrative access

Members are invited by email, assigned a role and workspace, and gain access after logging in with the invited address. All activity and credit usage is tracked per member, per workspace.

***

### Key Concepts

1. **Workspace Isolation** : Each member is assigned to a specific workspace. They can only see and use the workspace(s) assigned to them. This keeps project history clean and credit usage clearly attributed.
2. **Email-Matched Access :**  Invitations are tied to a specific email address. The invited user must log in — or create an account — using that exact email.
3. **Role-Controlled Governance :** Only Owners and Admins can invite members, modify roles, or manage workspaces. Members have no administrative access, ensuring a clear separation between usage and management.

***

### Who Should Read This

* **Owners** setting up their organization for the first time
* **Admins** managing a growing team
* **Anyone** troubleshooting access or onboarding a new team member


# Member Roles Explained

Learn how Owner, Admin, and Member roles control access, permissions, and responsibilities within an Organization.

Qolaba uses role-based access control (RBAC) to keep organizations structured, secure, and easy to manage. Every member of an organization holds one of three roles: **Owner**, **Admin**, or **Member**.

***

### Quick Comparison

<table><thead><tr><th width="228.0703125">Permission</th><th width="137.33984375">Owner</th><th width="140.8570556640625">Admin</th><th>Member</th></tr></thead><tbody><tr><td>Use AI tools (Chatbot, Image, Video, Speech)</td><td>✅</td><td>✅</td><td>✅</td></tr><tr><td>View Overview &#x26; History</td><td>✅</td><td>✅</td><td>✅ (assigned workspace)</td></tr><tr><td>Invite &#x26; remove members</td><td>✅</td><td>✅</td><td>❌</td></tr><tr><td>Assign &#x26; change roles</td><td>✅</td><td>✅</td><td>❌</td></tr><tr><td>Create, edit &#x26; delete workspaces</td><td>✅</td><td>✅</td><td>❌</td></tr><tr><td>Monitor credit usage</td><td>✅</td><td>✅</td><td>❌</td></tr><tr><td>Buy credits</td><td>✅</td><td>✅</td><td>❌</td></tr><tr><td>Upgrade or cancel subscription</td><td>✅</td><td>✅</td><td>❌</td></tr><tr><td>Edit organization profile</td><td>✅</td><td>✅</td><td>❌</td></tr><tr><td>Delete the organization</td><td>✅</td><td>❌</td><td>❌</td></tr></tbody></table>

***

### Owner

The Owner is the highest authority in an organization. The account that creates the organization is automatically assigned the Owner role.

**What the Owner controls:**

* All organization settings
* Subscription and billing
* Credit management
* Member and workspace management
* Structural deletion of the organization

**The one exclusive permission:** Only the Owner can permanently delete the organization. This restriction exists to prevent accidental or unauthorized removal of the entire account structure.

In most cases, an organization has a single Owner. The Owner role cannot be assigned through the standard invite flow — it transfers only through deliberate ownership change.

**Best suited for:** Founders, company principals, or whoever holds final accountability for the account.

***

### Admin

Admins support the Owner in day-to-day management. They have nearly full operational access, with one hard limit: they cannot delete the organization.

**What an Admin can do:**

* Invite and remove members
* Assign and change member roles (excluding ownership transfer)
* Create, edit, and delete workspaces
* Monitor credit usage across the organization
* Purchase additional credits
* Upgrade or cancel the subscription

**What an Admin cannot do:**

* Delete the organization

**Best suited for:** Team leads, operations managers, department heads, or project managers who need to manage people and workspaces without holding ultimate structural authority.

***

### Member

The Member role is designed for users who work within Qolaba — not those who manage it. Members have no administrative access.

**What a Member can do:**

* Access their assigned workspace
* Use all AI tools: Chatbot, Image generation, Video, Speech
* View the Overview tab
* View History within their assigned workspace

**What a Member cannot do:**

* Invite or remove members
* Create or manage workspaces
* Modify or cancel the subscription
* Purchase credits
* Edit organization settings
* Delete the organization

**Best suited for:** Designers, writers, marketers, developers, or any user whose primary need is to use Qolaba's AI tools.

***

### Why Role Separation Matters

Clear role boundaries prevent accidental changes, reduce governance risk, and make accountability straightforward as organizations scale.

**Example — Marketing Agency:**

<table><thead><tr><th width="296.6953125">Person</th><th>Role</th></tr></thead><tbody><tr><td>Founder</td><td>Owner</td></tr><tr><td>Account &#x26; Project Managers</td><td>Admin</td></tr><tr><td>Designers, Copywriters, Strategists</td><td>Member</td></tr></tbody></table>


# Inviting Members

Learn how to invite members, assign roles and workspaces, and manage access within your Organization.

Inviting a member adds them to your organization and grants them access to a specific workspace. The process takes under a minute — enter their email, assign a role, select a workspace, and send.

***

### Step-by-Step

1. Go to the **Members** tab in the Organization Dashboard.
2. Click **Invite Member**.
3. Enter the email address(es) of the person you want to invite.
   * You can enter a single address or multiple addresses at once.
   * Each address receives a separate invitation email.
4. Assign a **role** — Admin or Member. See [Member Roles Explained](/organization-and-workspaces/members-and-role-based-access/member-roles-explained) if you're unsure.
5. Select a **workspace** from the dropdown.
6. Send the invitation.

{% hint style="info" %}
The invited user must sign in using the exact email address they were invited with. Read more about [Member Activation Flow](/organization-and-workspaces/members-and-role-based-access/member-activation-flow)
{% endhint %}

***

### Workspace Assignment

Each member is assigned to exactly one workspace at the time of invitation. They will only have access to that workspace — nothing else in the organization is visible to them.

If your organization has multiple workspaces, double-check the assignment before sending. You can reassign a member to a different workspace later if needed.

***

### Managing Existing Members

Once a member is active, you can make the following changes from the **Members** tab:

1. **Change Role -** Promote a Member to Admin, or demote an Admin to Member. Role changes take effect immediately.
2. **Change Workspace -** Reassign a member to a different workspace when projects shift or teams reorganize.
3. **Revoke Access** - Removes the member from the organization immediately.&#x20;

***

### What Happens After You Send the Invitation

* The invited user receives an email notifying them they've been added to your organization and workspace.
* They log in to Qolaba with the invited email address.
* They switch to the organization from their profile dropdown.
* Their status updates to **Active** in the Members tab.
* They can now use AI tools, with all credit usage tracked individually.

For a full walkthrough of the member's side of this process, see [Member Activation Flow ](/organization-and-workspaces/members-and-role-based-access/member-activation-flow)


# Member Activation Flow

Learn how invited members activate their access and join their assigned workspace within an Organization.

After an Admin or Owner sends an invitation, the new member needs to complete a short activation process to access their workspace. This page walks through exactly what they need to do.

***

### Step 1 — Log In with the Invited Email

The invited user receives an invitation email. To activate their access:

1. Go to [Qolaba](https://qolaba.ai).
2. Log in using the **same email address** that received the invitation.

If they don't already have a Qolaba account, they must create one using that exact email address. Access will not be granted if the login email doesn't match the invited email — there are no workarounds for this.

***

### Step 2 — Switch to the Organization

After logging in, the user lands in their personal account by default. To access the organization they were invited to:

1. Click the **profile icon** in the top-right corner.
2. Open the profile dropdown.
3. Select the name of the organization.

The platform switches from personal mode to organization mode. From this point:

* Credits are drawn from the organization's credit pool, not the user's personal balance.
* All activity is recorded within the organization's account.
* Workspace permissions apply.

***

### Step 3 — Access the Assigned Workspace

Once in organization mode, the member is placed directly into their assigned workspace. They can confirm this by checking the workspace name in the top-left of the navigation panel.

Inside the workspace, the member can:

* Use Chatbot
* Generate images
* Create videos
* Use speech tools
* View history for that workspace

**What they can see:**

* Their assigned workspace(s)
* Credits available within the organization
* Their accessible dashboard tabs (Overview, History)

**What they cannot see:**

* Other workspaces within the organization
* Organization management settings
* Member management controls

***

### Troubleshooting

<table><thead><tr><th width="204.37188720703125">Issue</th><th>Likely Cause</th><th>Fix</th></tr></thead><tbody><tr><td>Invitation email not received</td><td>Email in spam, or wrong address entered</td><td>Check spam folder; ask Admin to re-invite with correct email</td></tr><tr><td>Access denied after login</td><td>Logged in with a different email</td><td>Log out and log in with the invited email address</td></tr><tr><td>Organization not appearing in dropdown</td><td>Account not yet activated</td><td>Ensure login email matches invited email exactly</td></tr><tr><td>Wrong workspace visible</td><td>Assigned to incorrect workspace</td><td>Ask Admin to reassign via the Members tab</td></tr></tbody></table>


# Tracking Credits and AI Usage

Monitor how AI credits are being consumed across your organization — broken down by workspace and by individual member.

You can access all usage data from **Dashboard → Members tab** or **Dashboard → Workspaces tab**.

***

#### Workspace-Level Usage

The **Workspaces tab** shows credit consumption for each workspace in your organization.

For each workspace, you can view credits used across:

<table><thead><tr><th width="207.90313720703125">Usage Type</th><th>What It Tracks</th></tr></thead><tbody><tr><td><strong>Chat</strong></td><td>Conversational AI interactions</td></tr><tr><td><strong>Image</strong></td><td>Image generation requests</td></tr><tr><td><strong>Speech</strong></td><td>Text-to-speech and voice output</td></tr><tr><td><strong>Total</strong></td><td>Combined usage across all types</td></tr></tbody></table>

Usage is shown relative to the workspace's **allocated budget** (in Budget Mode) or drawn from the **shared organizational pool** (in Shared Pool Mode). If a workspace is approaching or has exhausted its allocation, it will be reflected here.

{% hint style="info" %}
Use workspace-level tracking to identify which projects are consuming the most credits and rebalance allocations if needed. See [Workspace Credit Allocation](/organization-and-workspaces/workspaces/workspace-credit-allocation) for how to transfer credits between workspaces.
{% endhint %}

***

#### Member-Level Usage

The **Members tab** shows AI usage per individual across your organization.

For each member, you can view:

<table><thead><tr><th width="267.765625">Usage Type</th><th>What It Tracks</th></tr></thead><tbody><tr><td><strong>Chat</strong></td><td>AI chat sessions initiated by the member</td></tr><tr><td><strong>Image</strong></td><td>Image generation requests made by the member</td></tr><tr><td><strong>Speech</strong></td><td>Text-to-speech outputs generated by the member</td></tr><tr><td><strong>Total</strong></td><td>Sum of all usage types for that member</td></tr></tbody></table>

This helps you spot high-usage individuals, ensure fair credit distribution, and hold team members accountable for their consumption.

***

#### Quick Reference

<table><thead><tr><th width="280.1031494140625">Where to look</th><th>What you get</th></tr></thead><tbody><tr><td>Dashboard → Workspaces tab</td><td>Credit usage per workspace vs. allocated budget</td></tr><tr><td>Dashboard → Members tab</td><td>Credit usage per individual member</td></tr></tbody></table>


# Best Practices and Workflows

### 14.1 For Agencies

* Workspace per client
* Assign client-specific team members

### 14.2 For SMBs

* Workspace per department

### 14.3 For Consultants

* Workspace per engagement

### 14.4 For Freelancers

* Workspace per revenue stream

***

## Why This Structure Works

* Fully hierarchical
* Covers every tab and feature
* Separates role-based views
* Includes ICP use cases
* Scalable if features expand
* Layout-independent
* Enterprise-ready documentation


# Chatbot Models

A complete reference for all chatbot models available in Qolaba — context window, input and output credit costs, and best use cases organized by provider

Qolaba's Chatbot gives you access to 30+ large language models across six providers — all from a single interface. Each model has a different context window, credit cost, and area of strength. Use this page as a reference when selecting a model for a specific task.

***

### How to Read This Page

<table><thead><tr><th width="229.64453125">Column</th><th>What It Means</th></tr></thead><tbody><tr><td><strong>Context Window</strong></td><td>Maximum tokens the model can process in a single request — includes your prompt, conversation history, uploaded files, and the model's response. See <a href="/pages/0qEffBKJmqIjV3bW7CAg">Model Information Panel →</a> for a detailed explanation of how context windows work.</td></tr><tr><td><strong>Input Credits / 1K tokens</strong></td><td>Credits consumed per 1,000 input tokens — your prompt, files, conversation history, and system instructions</td></tr><tr><td><strong>Output Credits / 1K tokens</strong></td><td>Credits consumed per 1,000 output tokens — the model's generated response, including thinking tokens if Thinking Depth is enabled</td></tr></tbody></table>

{% hint style="info" %}
Models marked with ⭐ are available on **paid plans only**.
{% endhint %}

***

### Gemini Models

*Google*

Gemini models are Google's family of large language models — strong across long-context tasks, multimodal inputs, and general-purpose generation. The 1M token context window across most Gemini models makes them particularly well-suited for large document analysis, extended research sessions, and long conversations.

| Model                        | Context Window | Input Credits / 1K | Output Credits / 1K | Best For                                                                                |
| ---------------------------- | -------------- | ------------------ | ------------------- | --------------------------------------------------------------------------------------- |
| **Gemini 2.5 Flash**         | 1M             | 0.08               | 0.65                | Fast, cost-efficient general tasks — everyday queries, summarization, quick drafts      |
| **Gemini 2.5 Pro**           | 1M             | 0.33               | 0.39                | Balanced quality and cost — research, analysis, long-document processing                |
| **Gemini 3 Flash Preview**   | 1M             | 0.13               | 0.78                | Fast generation with improved quality over 2.5 Flash — content drafting, quick analysis |
| **Gemini 3 Pro Preview** ⭐   | 1M             | 0.52               | 3.12                | High-quality outputs — complex reasoning, detailed analysis, nuanced writing            |
| **Gemini 3.1 Pro Preview** ⭐ | 1M             | 1.04               | 4.68                | Highest quality Gemini output — advanced reasoning, complex multi-step tasks            |

***

### Claude Models

*Anthropic*

Claude models are Anthropic's family of large language models — known for strong instruction following, nuanced writing quality, and reliable performance on long-form content. Claude models have a 200K context window, making them well-suited for detailed documents, complex briefs, and extended reasoning tasks.

| Model                   | Context Window | Input Credits / 1K | Output Credits / 1K | Best For                                                                                |
| ----------------------- | -------------- | ------------------ | ------------------- | --------------------------------------------------------------------------------------- |
| **Claude Sonnet 4.6** ⭐ | 1M             | 0.78               | 3.90                | Balanced quality and speed — writing, analysis, coding, general professional tasks      |
| **Claude Opus 4.6** ⭐   | 200K           | 1.30               | 6.50                | Premium quality — complex reasoning, detailed writing, nuanced instruction following    |
| **Claude Opus 4.7** ⭐   | 200K           | 1.30               | 6.50                | Latest Opus — advanced reasoning, high-complexity tasks, long-form professional content |

***

### OpenAI Models

*OpenAI*

OpenAI models span a wide range — from the most cost-efficient nano models for everyday tasks to advanced reasoning models for complex problem solving. The GPT and o-series models offer strong prompt comprehension, reliable structured output, and broad capability across coding, writing, and analysis.

| Model              | Context Window | Input Credits / 1K | Output Credits / 1K | Best For                                                                                |
| ------------------ | -------------- | ------------------ | ------------------- | --------------------------------------------------------------------------------------- |
| **GPT-4.1** ⭐      | 1M             | 0.52               | 2.08                | General-purpose — reliable across writing, coding, analysis, and summarization          |
| **GPT-4.1 Mini**   | 1M             | 0.10               | 0.42                | Cost-efficient general tasks — everyday queries, drafts, quick summaries                |
| **GPT-5 Nano**     | 128K           | 0.03               | 0.10                | Most cost-effective OpenAI model — rapid iteration, high-volume simple tasks            |
| **GPT-5 Mini**     | 200K           | 0.12               | 0.94                | Lightweight everyday tasks — content drafting, quick answers, basic analysis            |
| **GPT-5.2** ⭐      | 200K           | 0.46               | 3.64                | Balanced quality — professional writing, structured analysis, coding assistance         |
| **GPT-5.2 Pro** ⭐  | 200K           | 5.46               | 43.68               | Maximum GPT-5.2 capability — highest quality structured outputs, complex reasoning      |
| **GPT-5.4** ⭐      | 272K           | 0.65               | 3.90                | Strong general capability — detailed analysis, complex writing, multi-step tasks        |
| **GPT-5.4 Mini**   | 200K           | 0.20               | 1.17                | Balanced speed and quality — content creation, moderate complexity tasks                |
| **GPT-5.4 Nano**   | 128K           | 0.05               | 0.33                | Fast, low-cost iteration — simple tasks, drafts, quick queries                          |
| **GPT-5.5** ⭐      | 1M             | 1.30               | 7.80                | Flagship GPT model — advanced reasoning, complex multi-step tasks, high-quality outputs |
| **OpenAI o1** ⭐    | 200K           | 4.29               | 17.16               | Advanced reasoning — complex logic, math, coding, multi-step problem solving            |
| **OpenAI o3** ⭐    | 200K           | 0.52               | 2.08                | Strong reasoning at moderate cost — analytical tasks, structured problem solving        |
| **OpenAI o4 Mini** | 200K           | 0.29               | 1.14                | Cost-efficient reasoning — logic tasks, coding, analysis at lower credit cost           |

***

### DeepSeek Models

*DeepSeek*

DeepSeek models deliver strong technical performance — particularly for coding, mathematical reasoning, and analytical tasks — at highly competitive credit costs. Well-suited for developer workflows and cost-sensitive high-volume use cases.

<table><thead><tr><th width="161.703125">Model</th><th>Context Window</th><th>Input Credits / 1K</th><th>Output Credits / 1K</th><th>Best For</th></tr></thead><tbody><tr><td><strong>DeepSeek V3.2</strong> ⭐</td><td>131K</td><td>0.07</td><td>0.10</td><td>Cost-efficient general tasks — coding assistance, technical writing, analysis</td></tr><tr><td><strong>DeepSeek V3.2 Speciale</strong> ⭐</td><td>128K</td><td>0.07</td><td>0.10</td><td>General-purpose tasks — everyday queries, content drafting, summarization</td></tr><tr><td><strong>DeepSeek R1</strong> ⭐</td><td>164K</td><td>0.18</td><td>0.65</td><td>Reasoning tasks — multi-step logic, math, structured problem solving</td></tr></tbody></table>

***

### Grok Models

*xAI*

Grok models are xAI's family of large language models — built for fast, real-time responses with strong general capability. The 2M token context window makes Grok models the highest context capacity models available in Qolaba — suited for extremely long documents, large codebases, and extended multi-turn conversations.

| Model               | Context Window | Input Credits / 1K | Output Credits / 1K | Best For                                                                                                     |
| ------------------- | -------------- | ------------------ | ------------------- | ------------------------------------------------------------------------------------------------------------ |
| **Grok 4.1 Fast** ⭐ | 2M             | 0.05               | 0.16                | Fast, cost-efficient responses — everyday tasks, quick analysis, real-time queries                           |
| **Grok 4.20** ⭐     | 2M             | 0.52               | 1.56                | High-quality responses with maximum context — large document analysis, extended conversations, complex tasks |

***

### Perplexity Models

*Sonar*

Perplexity's Sonar models are purpose-built for web-grounded responses — all models have built-in internet search, delivering answers backed by live, up-to-date sources rather than training data alone. Best for research, fact-checking, competitive intelligence, and any query where current information matters.

| Model                     | Context Window | Input Credits / 1K | Output Credits / 1K | Best For                                                                                         |
| ------------------------- | -------------- | ------------------ | ------------------- | ------------------------------------------------------------------------------------------------ |
| **Sonar**                 | 127K           | 0.26               | 0.26                | Fast web-grounded responses — general research, current events, quick fact-checking              |
| **Sonar Pro**             | 200K           | 0.78               | 3.90                | Higher quality web-grounded responses — detailed research, in-depth analysis with live sources   |
| **Sonar Reasoning Pro** ⭐ | 200K           | 0.52               | 2.08                | Web-grounded reasoning — research tasks requiring logical analysis of live information           |
| **Sonar Deep Research** ⭐ | 128K           | 0.52               | 2.08                | Deep, multi-source research — comprehensive reports, competitive analysis, thorough fact-finding |

***

### Choosing the Right Model

With 30+ models available, here is a practical starting point for common use cases:

| Use Case                                 | Recommended Model                   | Reason                                           |
| ---------------------------------------- | ----------------------------------- | ------------------------------------------------ |
| **Everyday tasks and drafting**          | Gemini 2.5 Flash or GPT-4.1 Mini    | Low cost, reliable quality for standard tasks    |
| **Professional writing and analysis**    | Claude Sonnet 4.6 or GPT-5.2        | Strong writing quality and instruction following |
| **Complex reasoning and logic**          | OpenAI o3 or DeepSeek R1            | Purpose-built for multi-step reasoning           |
| **Advanced reasoning — maximum quality** | OpenAI o1 or Gemini 3.1 Pro Preview | Highest reasoning capability available           |
| **Coding and technical tasks**           | DeepSeek V3.2 or GPT-5.4            | Strong technical performance at competitive cost |
| **Research with live web data**          | Sonar or Sonar Deep Research        | Built-in web search for current, sourced answers |
| **Long document analysis**               | Grok 4.20 or Gemini 2.5 Pro         | 1M–2M context window for large inputs            |
| **High-volume, cost-sensitive tasks**    | GPT-5 Nano or Grok 4.1 Fast         | Lowest credit cost per token                     |
| **Premium quality — best output**        | GPT-5.5 or Claude Opus 4.7          | Flagship models for highest quality output       |


# Image Models

A complete reference for all image generation models available in Qolaba — credit costs, quality options, generation limits, and best use cases organized by workflow type.

Qolaba provides access to a curated set of image generation models across text-to-image, image-to-image, and image editing workflows. Each model has distinct strengths, supported quality levels, credit costs, and generation limits. Use this page as a reference when selecting a model for your image generation task.

***

### How to Read This Page

| Column                   | What It Means                                                       |
| ------------------------ | ------------------------------------------------------------------- |
| **Credits / Image**      | Credits consumed per generated image at the specified quality level |
| **Quality**              | Supported resolution or quality tiers for that model                |
| **Max Generations**      | Maximum number of images producible in a single run                 |
| **Max Reference Images** | Maximum number of reference images uploadable (image-to-image only) |

{% hint style="info" %}
Models marked with ⭐ are available on **paid plans only**.
{% endhint %}

***

### Text-to-Image Models

Text-to-image models generate images from written prompts. Select based on the visual style, quality level, and credit budget required for your task.

#### **Flagship Models**

| Model                 | Quality | Credits / Image | Max Generations | Best For                                                                                   |
| --------------------- | ------- | --------------- | --------------- | ------------------------------------------------------------------------------------------ |
| **Nano Banana 2**     | 0.5K    | 15              | 4               | General-purpose generation — fast, versatile, strong quality across a wide range of styles |
|                       | 1K      | 21              |                 |                                                                                            |
|                       | 2K      | 32              |                 |                                                                                            |
|                       | 4K      | 48              |                 |                                                                                            |
| **Nano Banana Pro** ⭐ | 1K / 2K | 42              | 4               | Premium quality — higher fidelity for demanding, production-ready outputs                  |
|                       | 4K      | 75              |                 |                                                                                            |

***

#### **OpenAI Models**

GPT Image 2 uses quality tiers — **Low**, **Medium**, and **High** — instead of resolution values. These correspond to increasing levels of detail, prompt adherence, and visual fidelity:

* **Low** — fast generation, basic detail, suitable for drafts and concept testing
* **Medium** — balanced quality for most professional use cases
* **High** — maximum detail, photorealism, and prompt accuracy for final production output

| Model             | Quality | Credits / Image | Max Generations | Best For                                                                        |
| ----------------- | ------- | --------------- | --------------- | ------------------------------------------------------------------------------- |
| **GPT Image 2** ⭐ | Low     | 4               | 8               | Photorealism, text-in-image, structured scenes with strong prompt comprehension |
|                   | Medium  | 19              |                 |                                                                                 |
|                   | High    | 69              |                 |                                                                                 |

***

#### **Google Models**

| Model                  | Quality | Credits / Image | Max Generations | Best For                                                                           |
| ---------------------- | ------- | --------------- | --------------- | ---------------------------------------------------------------------------------- |
| **ImageGen 4**         | 1K / 2K | 11              | 4               | General-purpose — Google's latest image model, strong all-rounder for varied tasks |
| **ImageGen 4 Fast**    | 1K / 2K | 6               | 4               | Fast generation — quick iterations, draft testing, high-volume workflows           |
| **ImageGen 4 Ultra** ⭐ | 1K / 2K | 16              | 4               | Maximum quality — highest fidelity tier from Google for production-ready output    |

***

#### **Flux Models**

*Black Forest Labs*

| Model            | Quality  | Credits / Image | Max Generations | Best For                                                                                     |
| ---------------- | -------- | --------------- | --------------- | -------------------------------------------------------------------------------------------- |
| **Flux.2 Pro** ⭐ | Standard | 19              | 4               | Creative, high-fidelity generation — next-generation Flux model for detailed artistic output |
| **Flux.1 Dev**   | Standard | 8               | 4               | Developer-friendly experimentation — open model, great for testing and creative exploration  |

***

#### **ByteDance Models**

| Model            | Quality  | Credits / Image | Max Generations | Best For                                                                                             |
| ---------------- | -------- | --------------- | --------------- | ---------------------------------------------------------------------------------------------------- |
| **Seedream 4.5** | Standard | 13              | 6               | Creative and stylized images — ByteDance's flagship, exceptional for artistic and expressive outputs |

***

#### **Recraft Models**

| Model          | Quality  | Credits / Image | Max Generations | Best For                                                                                              |
| -------------- | -------- | --------------- | --------------- | ----------------------------------------------------------------------------------------------------- |
| **Recraft V4** | Standard | 13              | 4               | Design assets and icons — clean graphic design output, brand-consistent visuals, vector-style imagery |

***

### Image-to-Image Models

Image-to-image models accept one or more reference images to guide generation. Everything else — prompt, keywords, output controls — works the same as text-to-image. The reference images influence composition, style, subject, or structure of the output.

#### **Qolaba Flagship Models**

| Model                 | Quality | Credits / Image | Max Generations | Max Reference Images | Best For                                                                               |
| --------------------- | ------- | --------------- | --------------- | -------------------- | -------------------------------------------------------------------------------------- |
| **Nano Banana 2**     | 0.5K    | 15              | 4               | 13                   | General-purpose reference-guided generation — versatile, consistent quality            |
|                       | 1K      | 21              |                 |                      |                                                                                        |
|                       | 2K      | 32              |                 |                      |                                                                                        |
|                       | 4K      | 48              |                 |                      |                                                                                        |
| **Nano Banana Pro** ⭐ | 1K / 2K | 42              | 4               | 13                   | Premium reference-guided generation — highest fidelity for brand and production assets |
|                       | 4K      | 75              |                 |                      |                                                                                        |

***

#### **OpenAI Models**

<table><thead><tr><th width="113.29296875">Model</th><th>Quality</th><th>Credits / Image</th><th>Max Generations</th><th>Max Reference Images</th><th>Best For</th></tr></thead><tbody><tr><td><strong>GPT Image 2</strong> ⭐</td><td>Low</td><td>4</td><td>8</td><td>15</td><td>Photorealistic transformations, text-in-image, strong prompt-guided reference editing</td></tr><tr><td></td><td>Medium</td><td>19</td><td></td><td></td><td></td></tr><tr><td></td><td>High</td><td>69</td><td></td><td></td><td></td></tr></tbody></table>

***

#### **Flux Models**

| Model            | Quality  | Credits / Image | Max Generations | Max Reference Images | Best For                                                                         |
| ---------------- | -------- | --------------- | --------------- | -------------------- | -------------------------------------------------------------------------------- |
| **Flux.2 Pro** ⭐ | Standard | 19              | 4               | 8                    | Creative, high-fidelity style and content transformation with reference guidance |

***

### Image Editing Models

Image editing models are purpose-built tools used within the four editing workflows — Inpainting, Background Removal, Upscaling, and Image Variation. They run automatically based on the tool selected and are not manually chosen in the model selector.

| Model                       | Used For             | What It Does                                                                                        |
| --------------------------- | -------------------- | --------------------------------------------------------------------------------------------------- |
| **BiRefNet V2**             | Background Removal   | One-click background removal — accurately detects and isolates the subject from any background      |
| **Flux General Inpainting** | Inpainting & Cleanup | Mask-based editing — removes or replaces specific areas of an image based on a text prompt          |
| **Nano Banana 2**           | Image Variation      | Generates creative variations inspired by uploaded reference images — up to 13 references supported |

***

### Model Comparison at a Glance

| Use Case                           | Recommended Model                  | Reason                                                                           |
| ---------------------------------- | ---------------------------------- | -------------------------------------------------------------------------------- |
| **General-purpose generation**     | Nano Banana 2                      | Versatile, strong quality, lowest cost entry point for flagship output           |
| **Premium production output**      | Nano Banana Pro ⭐                  | Highest fidelity for demanding, client-facing creative work                      |
| **Photorealism and text-in-image** | GPT Image 2 ⭐                      | OpenAI's latest — best-in-class for photorealistic scenes and legible text       |
| **Creative and artistic output**   | Seedream 4.5                       | Exceptional for stylized, expressive, and artistic image generation              |
| **Design assets and icons**        | Recraft V4                         | Clean graphic design output with strong brand consistency                        |
| **Fast iteration and drafts**      | ImageGen 4 Fast or Flux.1 Dev      | Low cost, fast generation — ideal for prompt testing and concept exploration     |
| **Maximum Google quality**         | ImageGen 4 Ultra ⭐                 | Google's highest quality tier for production-ready output                        |
| **Reference-guided generation**    | Nano Banana 2 or Nano Banana Pro ⭐ | Up to 13 reference images — best for brand consistency and style matching        |
| **Most reference images**          | GPT Image 2 ⭐                      | Supports up to 15 reference images per generation                                |
| **Background removal**             | BiRefNet V2 (auto)                 | Best-in-class subject isolation — runs automatically via Background Removal tool |
| **Inpainting and cleanup**         | Flux General Inpainting (auto)     | Precise mask-based editing — runs automatically via Inpainting tool              |


# Video Models

A complete reference for all video generation models available in Qolaba — credit costs by duration and resolution, supported features, limitations, and best use cases.

Qolaba provides access to 10+ video generation models — covering text-to-video, image-to-video, multi-reference guided generation, and AI-powered video editing. Credit costs vary by model, duration, and resolution. Use this page as a reference when selecting a model for your video generation task.

***

### How to Read This Page

<table><thead><tr><th width="222.83203125">Column</th><th>What It Means</th></tr></thead><tbody><tr><td><strong>Credits</strong></td><td>Credits consumed per video at the specified duration and resolution combination</td></tr><tr><td><strong>Duration</strong></td><td>Supported video lengths in seconds</td></tr><tr><td><strong>Resolution</strong></td><td>Supported output quality options</td></tr><tr><td><strong>Reference Images</strong></td><td>Whether the model accepts uploaded images as generation input</td></tr><tr><td><strong>Audio Support</strong></td><td>Whether the model supports AI-generated or uploaded audio</td></tr></tbody></table>

{% hint style="info" %}
Models marked with ⭐ are available on **paid plans only**.
{% endhint %}

***

### **Google Veo Models**

Google's flagship video generation models — delivering cinematic quality, physics-accurate motion, and highly detailed environments.

#### **Veo 3.1** ⭐

| Duration | 720p / 1080p | 4K    |
| -------- | ------------ | ----- |
| 4s       | 416          | —     |
| 6s       | 624          | —     |
| 8s       | 832          | 1,248 |

#### **Veo 3.1 Fast**

| Duration | 720p / 1080p | 4K  |
| -------- | ------------ | --- |
| 4s       | 156          | —   |
| 6s       | 234          | —   |
| 8s       | 312          | 728 |

| Feature                | Veo 3.1                                                           | Veo 3.1 Fast                                                 |
| ---------------------- | ----------------------------------------------------------------- | ------------------------------------------------------------ |
| **Input**              | Text-to-video, Image-to-video                                     | Text-to-video, Image-to-video                                |
| **Default resolution** | 720p                                                              | 720p                                                         |
| **Max generations**    | 4                                                                 | 4                                                            |
| **Best for**           | Cinematic quality, physics-accurate motion, detailed environments | Faster generation at lower cost — balanced speed and quality |

**Duration & Resolution Restrictions:**

| Condition                                 | Allowed Durations |
| ----------------------------------------- | ----------------- |
| 720p + text-to-video (no reference image) | 4s, 6s, 8s        |
| 720p + reference image (image-to-video)   | 8s only           |
| 1080p (any input)                         | 8s only           |
| 4K (any input)                            | 8s only           |

{% hint style="info" %}
4K is only available at 8 seconds duration for both Veo models.
{% endhint %}

***

### **Runway Models**

#### **Runway Gen-4.5** ⭐

Credits are calculated at 31.2 credits per second:

| Duration | Credits |
| -------- | ------- |
| 2s       | 63      |
| 5s       | 156     |
| 10s      | 312     |

| Feature                 | Details                                                                           |
| ----------------------- | --------------------------------------------------------------------------------- |
| **Input**               | Text-to-video, Image-to-video                                                     |
| **Supported durations** | 2–10 seconds                                                                      |
| **Output resolution**   | 720p only                                                                         |
| **Frame rate**          | 24fps, 25fps                                                                      |
| **Max generations**     | 4                                                                                 |
| **Best for**            | Industry-leading production quality — reliable for commercial and branded content |

**Aspect Ratio Restrictions:**

| Input Mode     | Supported Aspect Ratios         |
| -------------- | ------------------------------- |
| Text-to-video  | 16:9 only                       |
| Image-to-video | 16:9, 9:16, 1:1, 4:3, 3:4, 21:9 |

***

### **ByteDance Models**

**Seedance 2.0** and **Seedance 2.0 Fast** are ByteDance's flagship video models — distinguished by their multi-reference input capability, allowing up to 12 reference files (images, videos, and audio) to guide a single generation.

***

#### **Seedance 2.0** ⭐

| Feature                   | Details                                                                                                          |
| ------------------------- | ---------------------------------------------------------------------------------------------------------------- |
| **Input**                 | Text-to-video, Image-to-video, Video-to-video                                                                    |
| **Supported durations**   | 4–15 seconds (or AI-determined if left blank)                                                                    |
| **Supported resolutions** | 480p, 720p, 1080p                                                                                                |
| **Aspect ratios**         | 16:9, 21:9, 9:16, 3:4, 1:1, 4:3                                                                                  |
| **Max generations**       | 4                                                                                                                |
| **Audio support**         | AI-generated audio or uploaded audio                                                                             |
| **Best for**              | Cinematic quality with multi-reference guidance — brand-consistent generation combining images, video, and audio |

#### **Seedance 2.0 Fast** ⭐

| Feature                   | Details                                                                  |
| ------------------------- | ------------------------------------------------------------------------ |
| **Input**                 | Text-to-video, Image-to-video, Video-to-video                            |
| **Supported durations**   | 4–15 seconds                                                             |
| **Supported resolutions** | 480p, 720p                                                               |
| **Max generations**       | 4                                                                        |
| **Audio support**         | AI-generated audio or uploaded audio                                     |
| **Best for**              | Faster generation at lower cost — rapid prototyping and quick iterations |

**Reference Media Support (Both Seedance Models):**

| Media Type      | Limit                   | Size Limit                              |
| --------------- | ----------------------- | --------------------------------------- |
| Images          | Up to 9                 | Max 30 MB each                          |
| Videos          | Up to 3                 | Combined 2–15 seconds, max 50 MB total  |
| Audio           | Up to 3                 | Combined max 15 seconds, max 15 MB each |
| **Total files** | Max 12 across all types | —                                       |

**How to Reference Uploads in Your Prompt:**

Use tags to tell the model exactly how to use each uploaded file:

* Images: `@Image1`, `@Image2`, etc.
* Videos: `@Video1`, `@Video2`, etc.
* Audio: `@Audio1`, `@Audio2`, etc.

**Example prompt:**

```
The person in @Image1 walks into a futuristic city
while @Audio1 plays softly in the background.
```

> **Note:** Audio cannot be uploaded without at least one image or video reference. Maximum 12 files total across all media types.

***

### **Happy Horse Models**&#x20;

Happy Horse models are purpose-built for strong character consistency and advanced video editing — the only models in Qolaba that support AI-powered editing of existing video footage.

***

#### **Happy Horse (Generation)** ⭐

| Feature                   | Details                                                                           |
| ------------------------- | --------------------------------------------------------------------------------- |
| **Input**                 | Text-to-video, Image-to-video                                                     |
| **Supported durations**   | 3–15 seconds (default 5s)                                                         |
| **Supported resolutions** | 720p, 1080p (default 1080p)                                                       |
| **Aspect ratios**         | 16:9, 9:16, 1:1, 4:3, 3:4                                                         |
| **Max generations**       | 4                                                                                 |
| **Max reference images**  | Up to 9 images                                                                    |
| **Best for**              | Multi-character video generation with consistent character identity across scenes |

**How Character References Work:** Upload images of your characters — the first uploaded image becomes `character1`, the second becomes `character2`, and so on up to `character9`. Reference them directly in your prompt:

```
A futuristic dance battle between character1 and character2
under neon lights.
```

***

#### **Happy Horse Video Edit** ⭐

Happy Horse Video Edit is a distinct capability from generation — it edits and transforms existing videos rather than creating new ones from scratch.

| Feature                       | Details                                                                           |
| ----------------------------- | --------------------------------------------------------------------------------- |
| **Input**                     | Source video (required) + optional style images                                   |
| **Supported resolutions**     | 720p, 1080p                                                                       |
| **Source video requirements** | MP4 or MOV, 3–60 seconds, under 100 MB, minimum 320px shortest side               |
| **Max style images**          | Up to 5                                                                           |
| **Audio options**             | Keep original audio or regenerate with AI                                         |
| **Max output length**         | 15 seconds (longer videos are automatically trimmed)                              |
| **Best for**                  | Style transfer, element replacement, re-texturing or re-lighting existing footage |

**Example prompt:**

```
Replace the person's jacket with the red leather jacket
shown in @Image1, and make the sky look like a sunset.
```

> **Important:** Even if you upload a 60-second source video, the model processes and returns a maximum of 15 seconds of edited footage.

***

### **Minimax Models**

#### **Hailuo 2.3 Pro** ⭐

| Duration | Resolution | Credits |
| -------- | ---------- | ------- |
| 6s       | 1080p      | 128     |
| 10s      | 768p       | —       |

| Feature              | Details                                                                         |
| -------------------- | ------------------------------------------------------------------------------- |
| **Input**            | Text-to-video, Image-to-video                                                   |
| **Default duration** | 6 seconds                                                                       |
| **Max generations**  | 4                                                                               |
| **Best for**         | Strong general-purpose video generation — reliable motion and scene consistency |

> **Note:** 10-second duration is only available at 768p resolution.

***

### **Kling Models**

*Kuaishou*

#### **Kling O3 Pro** ⭐

| Duration | Credits |
| -------- | ------- |
| 5s       | 183     |
| 10s      | 365     |
| 15s      | 547     |

#### **Kling V3 Pro** ⭐

| Duration | Credits |
| -------- | ------- |
| 5s       | 219     |
| 10s      | 437     |
| 15s      | 656     |

| Feature                 | Kling O3 Pro                                                                        | Kling V3 Pro                                                        |
| ----------------------- | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------- |
| **Input**               | Text-to-video, Image-to-video                                                       | Text-to-video, Image-to-video                                       |
| **Supported durations** | 3–15 seconds                                                                        | 3–15 seconds                                                        |
| **Resolution**          | 720p / 1080p                                                                        | 720p / 1080p                                                        |
| **Max generations**     | 4                                                                                   | 4                                                                   |
| **Best for**            | High quality, strong motion — reliable for short-form social and commercial content | Latest Kling generation — improved motion realism and visual detail |

***

### **Vidu Models**

#### **Vidu Q3 Turbo**

| Duration | 360p / 540p | 720p / 1080p |
| -------- | ----------- | ------------ |
| 4s       | 37          | 81           |
| 8s       | 73          | 161          |
| 16s      | 146         | 321          |

| Feature                 | Details                                                                                                    |
| ----------------------- | ---------------------------------------------------------------------------------------------------------- |
| **Input**               | Text-to-video, Image-to-video                                                                              |
| **Supported durations** | 1–16 seconds                                                                                               |
| **Max generations**     | 4                                                                                                          |
| **Best for**            | Fast and cost-effective generation — quick cuts, looping visuals, short transitions, high-volume iteration |

***

### **Luma Models**

#### **Luma Ray 2**

| Duration | 540p | 720p | 1080p |
| -------- | ---- | ---- | ----- |
| 5s       | 130  | 260  | 520   |
| 9s       | 234  | 468  | 936   |

| Feature             | Details                                                                                             |
| ------------------- | --------------------------------------------------------------------------------------------------- |
| **Input**           | Text-to-video, Image-to-video                                                                       |
| **Max generations** | 4                                                                                                   |
| **Best for**        | Cinematic quality with excellent prompt adherence — smooth motion and strong narrative storytelling |

***

### **xAI Models**

#### **Grok Imagine Video** ⭐

| Duration | 480p | 720p |
| -------- | ---- | ---- |
| 5s       | 65   | 92   |
| 10s      | 130  | 183  |
| 15s      | 195  | 274  |

| Feature                 | Details                                                                                       |
| ----------------------- | --------------------------------------------------------------------------------------------- |
| **Input**               | Text-to-video, Image-to-video                                                                 |
| **Supported durations** | 1–15 seconds                                                                                  |
| **Max generations**     | 4                                                                                             |
| **Best for**            | Cost-effective generation across extended durations — creative and experimental video content |

***

### Model Comparison at a Glance

| Use Case                                 | Recommended Model                 | Reason                                                               |
| ---------------------------------------- | --------------------------------- | -------------------------------------------------------------------- |
| **Cinematic quality — maximum fidelity** | Veo 3.1 or Seedance 2.0 ⭐         | Google's flagship or ByteDance's premium model                       |
| **Fast generation — lower cost**         | Veo 3.1 Fast or Seedance 2.0 Fast | Faster variants at reduced credit cost                               |
| **Multi-reference guided generation**    | Seedance 2.0 ⭐                    | Only model supporting up to 12 reference files (image, video, audio) |
| **Character-consistent generation**      | Happy Horse ⭐                     | Dedicated character reference system for multi-character scenes      |
| **AI video editing**                     | Happy Horse Video Edit ⭐          | Only model supporting AI-powered editing of existing video footage   |
| **Industry-leading production quality**  | Runway Gen-4.5 ⭐                  | Reliable for commercial and branded content                          |
| **Long-duration generation**             | Kling O3 Pro or Kling V3 Pro ⭐    | Supports up to 15 seconds                                            |
| **Highest context window**               | Grok Imagine Video                | Cost-effective across 1–15 second durations                          |
| **Fast, low-cost iteration**             | Vidu Q3 Turbo                     | Lowest credit cost — good for drafts and quick concepts              |
| **Smooth cinematic motion**              | Luma Ray 2                        | Excellent prompt adherence with cinematic quality                    |
| **Strong general-purpose**               | Hailuo 2.3 Pro                    | Reliable motion and scene consistency                                |


# Speech Models

A complete reference for speech generation models in Qolaba — Gemini Flash TTS and Pro TTS, available voices, accent options, and best use cases.

Qolaba's Speech Generation is powered by **Google's Gemini TTS** — two models offering different trade-offs between speed, quality, and credit cost. Both models support single-speaker narration and multi-speaker dialogue generation with customizable voice, accent, and delivery style.

***

### Available Models

| Model                | Speed           | Quality                              | Best For                                                                 |
| -------------------- | --------------- | ------------------------------------ | ------------------------------------------------------------------------ |
| **Gemini Flash TTS** | Fast            | Good                                 | Script drafting, quick iterations, testing voice and accent combinations |
| **Gemini Pro TTS** ⭐ | Slightly slower | Higher — more expressive and natural | Final production output, client-ready audio, publishing                  |

{% hint style="info" %}
Test your script, voice selection, and accent configuration with Flash TTS first. Switch to Pro TTS for the final generation only. This approach saves credits without compromising final output quality.
{% endhint %}

***

### Model Capabilities

Both models share the same core capabilities:

| Capability                     | Flash TTS        | Pro TTS          |
| ------------------------------ | ---------------- | ---------------- |
| **Single Speaker Mode**        | ✓                | ✓                |
| **Multi-Speaker Mode**         | ✓                | ✓                |
| **Voice library**              | 30+ voices       | 30+ voices       |
| **Accent & dialect selection** | ✓                | ✓                |
| **Style instructions**         | ✓                | ✓                |
| **Multi-language support**     | ✓                | ✓                |
| **Max script length**          | 5,000 characters | 5,000 characters |

***

### Voice Library

Both models provide access to a library of 30+ distinct voice profiles. Each voice has a unique combination of tone, pitch, energy, and speaking style. Click any voice in the interface to hear a preview before selecting.

**Available Voices**

| Voice             | Character                      |
| ----------------- | ------------------------------ |
| **Zephyr**        | Bright, clear, and energetic   |
| **Puck**          | Upbeat and playful             |
| **Charon**        | Informative and measured       |
| **Kore**          | Firm and confident             |
| **Fenrir**        | Excitable and expressive       |
| **Aoede**         | Smooth and warm                |
| **Leda**          | Youthful and approachable      |
| **Orus**          | Clear and neutral              |
| **Schedar**       | Gravelly and deep              |
| **Gacrux**        | Soft and gentle                |
| **Pulcherrima**   | Calm and composed              |
| **Achird**        | Conversational and natural     |
| **Zubenelgenubi** | Professional and authoritative |
| **Vindemiatrix**  | Storytelling and expressive    |
| **Sadachbia**     | Warm and personable            |
| **Sadaltager**    | Crisp and articulate           |
| **Sulafat**       | Rich and resonant              |

{% hint style="info" %}
The full voice library is accessible directly in the [Speech Generation workspace](https://www.qolaba.ai/ai-speech-generator/text-to-speech). Click any voice name to preview it before selecting.
{% endhint %}

***

### Accent & Dialect Support

The output language is determined automatically by the language of your input script — write in any language and the audio is generated in that language. Accent selection refines pronunciation for languages with multiple regional dialects.

#### **Available Accents by Language**

| Language       | Available Dialects                              |
| -------------- | ----------------------------------------------- |
| **English**    | United States, United Kingdom, India, Australia |
| **French**     | France, Canada                                  |
| **Spanish**    | Spain, Latin America                            |
| **Arabic**     | Egypt, Global                                   |
| **Mandarin**   | China, Taiwan                                   |
| **Hindi**      | India                                           |
| **Portuguese** | Brazil, Portugal                                |
| **German**     | Germany, Austria                                |
| **Japanese**   | Japan                                           |
| **Korean**     | Korea                                           |

{% hint style="info" %}
The output language always follows the language of your input text. Accent selection narrows the regional dialect within that language — it does not override the language itself.
{% endhint %}

***

### Style Instructions

Both models support style instructions — a plain-language description of the desired delivery tone and manner entered in the **Style Prompt** field.

**Examples:**

<table><thead><tr><th width="273.52734375">Intent</th><th>Style Instruction</th></tr></thead><tbody><tr><td>Warm and conversational</td><td><em>"Speak warmly and conversationally, like talking to a friend"</em></td></tr><tr><td>Professional narration</td><td><em>"Clear, professional, and authoritative tone"</em></td></tr><tr><td>Energetic marketing</td><td><em>"Enthusiastic and high-energy delivery"</em></td></tr><tr><td>Calm instructional</td><td><em>"Calm, slow-paced, and easy to follow"</em></td></tr><tr><td>Storytelling</td><td><em>"Engaging narrative style with natural pauses and expression"</em></td></tr></tbody></table>

***

### Flash TTS vs. Pro TTS — When to Use Each

#### **Use Flash TTS when:**

* Testing a new script for the first time
* Validating voice, accent, and style combinations before final generation
* Producing audio for internal use, drafts, or non-published content
* Working at high volume where credit efficiency matters

#### **Use Pro TTS when:**

* Generating final production audio for publishing
* Delivering client-ready voiceovers, ads, or podcast content
* The naturalness and expressiveness of the voice matters for the audience
* Multi-speaker dialogue needs to sound as realistic as possible


# Music Models

A detailed breakdown of all Music Generation models in Qolaba

Two AI models are available for Music Generation in Qolaba — each with different output length, supported formats, configuration options, and credit costs. Choosing the right model depends on whether you need a short clip or a full-length song, and how much control you want over the output structure.

***

#### Model Overview

| Model             | Provider | Duration     | Credits | Format   | Lyrics |
| ----------------- | -------- | ------------ | ------- | -------- | ------ |
| **Lyria 3 Clip**  | Google   | \~30 seconds | 11      | MP3      | ✓      |
| **Lyria 3 Pro** ⭐ | Google   | Full-length  | 21      | MP3, WAV | ✓      |

***

#### Lyria 3 Clip

Lyria 3 Clip is Google's short-form music generation model — producing approximately 30-second high-fidelity clips from text or image prompts. It is the only model available to all users on free and paid plans.

**Capabilities:**

* Text and image prompt support
* Negative prompt support
* Seed support for reproducible output
* Lyrics generation

**Best for:**

* Social media clips and short-form content
* Background audio for videos and presentations
* Quick concept testing before committing to a full-length track
* Users on free plans who need music generation access

***

#### Lyria 3 Pro ⭐

Lyria 3 Pro is Google's full-length music generation model — producing multi-minute songs with complete structure including verses, choruses, and bridges. It is the most capable model for production-ready music output.

**Capabilities:**

* Text and image prompt support
* Negative prompt support
* Seed support for reproducible output
* Lyrics generation with time-sync support
* WAV format support for high-fidelity editing

**Best for:**

* Full-length songs for publishing, streaming, or client delivery
* Production-ready audio requiring WAV format
* Complete narrative songs with structured verses and choruses
* Professional music content creation

***

#### Configuration Options

**Instrumental Only**

Toggle on to generate music without vocals. Available across all models — useful for background tracks, soundscapes, and any content where vocals are not needed.

**Format Selection**

| Format  | File Size | Best For                                       |
| ------- | --------- | ---------------------------------------------- |
| **MP3** | Smaller   | Sharing, social media, web use                 |
| **WAV** | Larger    | High-fidelity editing, professional production |

{% hint style="info" %}
WAV is only available on Lyria 3 Pro. All other models output MP3 only.
{% endhint %}

**Negative Prompt**

Describe what to exclude from the generation — instruments, styles, tempos, or moods you do not want.&#x20;

**Examples:**

```
No drums, no distortion, no fast tempo.
```

```
Avoid brass instruments, no vocals, no reverb.
```

**Seed**

Enter a number to make your generation reproducible. Using the same prompt, model, and seed number produces the same output every time — useful for iterating on a specific result or sharing a configuration with a collaborator.

**Magic Prompt**

One-click AI enhancement that rewrites your prompt into a more detailed, model-optimized version before generation.

* **Cost:** \~0.05 credits
* **Best for:** Short or underdeveloped prompts

***

#### Choosing the Right Model

| Need                     | Recommended Model             |
| ------------------------ | ----------------------------- |
| Free plan access         | Lyria 3 Clip                  |
| Quick 30-second clip     | Lyria 3 Clip                  |
| Full-length song         | Lyria 3 Pro ⭐                 |
| WAV format output        | Lyria 3 Pro ⭐                 |
| Reference image guidance | Lyria 3 Clip or Lyria 3 Pro ⭐ |
| Most cost-efficient      | Lyria 3 Clip (11 credits)     |


# Introduction

An overview of the Qolaba Chatbot — what it is, who it's for, and a summary of everything it can do.

Qolaba's Chatbot is a multi-model AI assistant built for professional and team-based workflows. It goes beyond standard chat by combining access to multiple large language models, executable tools, custom agents, and knowledge-aware responses — all within your organization's workspace structure

### Who It's For

The Qolaba Chatbot is built for teams and individuals who need more than a basic AI chat interface:

* **Startups & Creators** — Rapid content generation, product documentation, growth experimentation
* **Marketing Teams** — Campaign ideation, ad copy, SEO drafts
* **Agencies** — Multi-client workspace management with shared credit pools
* **Developers** — Prompt engineering, model comparison, AI experimentation
* **Enterprises** — Structured AI collaboration with governance and audit trails

***

### Key Capabilities

Five capabilities make the Qolaba Chatbot distinct from a standard AI chat tool.

#### **1. Multi-LLM Access**

Most AI tools lock you into a single model. Qolaba gives you access to six leading large language models from one interface — **GPT, Gemini, Claude, DeepSeek, Grok, and Sonar (Perplexity)** — and lets you switch between them freely.

This means you can:

* Run the same prompt across multiple models and compare outputs
* Choose the best model for a specific task — creative writing, coding, research, summarization
* Optimize credit usage by matching task complexity to the right model
* Avoid being locked into one provider's strengths or limitations\
  [Learn more about Model Selection →](/model-reference/chatbot-models)

***

#### **2. Agents**

Agents are purpose-built AI assistants that go beyond a one-off chat. Each agent has a defined **role, behavioral instructions, tone, and knowledge base** — so it responds consistently, stays on-topic, and understands your context without you re-explaining it every session.

Qolaba offers two types:

* **Pre-built agents** — Ready-to-use assistants for common roles like Marketing Strategist, Content Creator, and more. Select one and start immediately with no setup required.
* **Custom agents** — Build your own from scratch. Define the agent's name, role, model, expertise, and attach knowledge bases so it answers with your specific context in mind.

Agents are persistent, reusable across sessions, and can be shared within a workspace — making them especially valuable for teams with recurring workflows.

[Learn more about Agents →](/chatbot/agents)

***

#### **3. Toolkit**

The Toolkit turns the Chatbot from a conversational tool into an execution environment. Instead of just generating text, it can actively perform tasks — searching the web, generating media, running code, analyzing files, and more.

Tools can run in **Auto mode**, where Qolaba automatically selects the right tool based on your prompt, or you can enable them manually for full control.

Toolkit includes tools for web search, media generation, code execution, file and URL analysis, and output safety — covering most tasks a professional workflow demands, without leaving the chat interface.

[Learn more about the Toolkit →](/chatbot/toolkit)

***

#### **4. Knowledge Base & Files**

Uploading knowledge base and files lets you extend the chatbot's context with your own material — so it can answer questions, summarize content, and extract insights from documents that are specific to your work.

You can:

* **Upload files directly** into a chat — supported formats include PDF, CSV, Excel, and plain text
* **Create knowledge bases** — organized collections of files that any agent or chat session can reference persistently

This transforms the Chatbot from a generic assistant into a domain-aware one that understands your products, processes, and data.

[Learn more about Knowledge Bases & Files →](/chatbot/adding-context-and-resources)

***

#### **5. Chat Branching**

When a response isn't quite right, you don't need to start over or manually re-prompt. Chat Branching lets you **regenerate any response using a different model** at any point in a conversation — and compare outputs side-by-side across LLMs.

This is particularly useful for:

* Testing how different models interpret the same prompt
* Selecting the best response before using it in production
* Iterating on complex outputs like code, copy, or structured data without losing conversation context

[Learn more about Chat Branching →](/chatbot/model-selection/chat-branching)

***

#### Additional Features

The Chatbot also includes several supporting features covered in their own pages:

* [**Model Settings**](/chatbot/model-selection/model-settings) — Fine-tune Temperature and Thinking Depth to control response style and reasoning
* [**Chat Management**](/chatbot/chat-history-management) — Pin, rename, share, and delete conversations from your history
* [**Prompt Management**](/chatbot/prompt-and-controls) — Save, edit, and reuse prompts across chats and agents

***

#### What's Next

Follow the recommended flow to get started:

1. [Start a New Chat →](/chatbot/starting-a-new-chat)
2. [Select a model and configure settings →](/chatbot/model-selection)
3. [Set up an agent →](/chatbot/agents)
4. [Upload files or create a knowledge base →](/chatbot/adding-context-and-resources)
5. [Explore the Toolkit →](/chatbot/toolkit)


# Starting a New Chat

How to start a new chat in the Qolaba Chatbot, when to do it, and why managing context matters.

Each chat in Qolaba is an isolated conversation thread tied to your current workspace. Starting a new chat gives you a clean context, a fresh message thread, and no influence from previous conversations.

#### Creating a New Chat

1. Open your **Workspace** from the left navigation panel
2. Select **Chatbot**
3. Click **New Chat** in the top right
4. Enter your first prompt and send

A new conversation thread is created automatically, saved to your workspace history, and begins consuming credits from your first message.

***

#### How Conversation Context Works

Every chat maintains its own **context window** — a memory of everything said within that thread, including your prompts, AI responses, follow-up instructions, and any files attached during the session.

This allows the AI to:

* Build on earlier instructions without you repeating them
* Refine outputs progressively across multiple turns
* Handle complex, multi-step tasks within a single conversation

{% hint style="info" %}
Context is chat-specific. Starting a new chat resets conversational memory entirely. AI has no memory of previous chat sessions.
{% endhint %}

***

#### When to Start a New Chat

Start a new chat when you are:

* Switching to a different topic or task
* Working on a new client or project
* Deliberately isolating results from a previous session

{% hint style="info" %}
**Best practice:** One chat per objective. It keeps your history clean, your outputs traceable, and your context uncontaminated.
{% endhint %}

***

#### Resetting Context

The cleanest way to reset context is to **start a new chat**. This ensures the AI has no memory of earlier instructions, files, or outputs — giving you a reliable baseline for experimentation or benchmarking.

Carrying over unintended context can cause the AI to reference outdated instructions, produce inconsistent comparisons, or generate outputs biased by earlier constraints.


# Chat History Management

How to manage your chat history in Qolaba using pin, rename, share, and delete actions.

Every conversation in Qolaba is automatically saved to your workspace history. The three-dot menu **(⋮)** on any chat thread gives you full control over how conversations are organized, shared, and maintained — keeping your workspace clean and easy to navigate.

***

### Viewing Chat History

All conversations are saved automatically in **reverse chronological order** — most recent at the top. Pinned chats remain fixed at the top of the list regardless of newer activity below them.

***

### Searching Chat History

Use the **search bar** in the chat history panel to quickly locate any past conversation. You can search by chat title, keywords from your prompts, or keywords from AI responses. Results filter automatically as you type.

***

### Managing Chats

Hover over any chat in your history panel and click the **⋮ menu** to the right of the chat title to access all management actions.

***

#### **Pin Chat**

Pinning keeps a conversation fixed at the top of your history list, regardless of how many new chats are created after it. Use this for active projects, ongoing client work, threads with finalized outputs, or conversations containing prompts you return to regularly.

{% hint style="info" %}
**Best practice:** Pin only active or high-priority threads. Over-pinning reduces the benefit and clutters your history panel.
{% endhint %}

***

#### **Rename Chat**

Chat titles are auto-generated from your first prompt — which is rarely descriptive enough for a growing history. Renaming makes your conversations scannable, searchable, and easier to share with teammates.

**To rename a chat:**

1. Click the **⋮ menu** on the chat
2. Select **Rename Chat**
3. Enter a clear, descriptive title
4. Save

**Naming conventions that work well:**

* `Client A – Blog Draft`
* `Product Launch – Ad Copy`
* `GPT vs Claude – Benchmark Test`
* `Internal Strategy – Q2 Planning`

Good naming improves searchability, workspace organization, and collaboration — especially in shared workspaces with multiple team members.

***

#### **Share Chat**

Share Chat generates a **read-only link** to a snapshot of the conversation. This is useful for client approvals, internal reviews, handoffs, or documentation.

**To generate a shareable link:**

1. Click the **⋮ menu** on the chat
2. Select **Share Chat**
3. Copy the generated link and distribute it

The recipient can view the conversation content but cannot edit it, access workspace settings, or interact with your credits. Only the conversation itself is shared.

**To revoke access:**

1. Return to **Share Chat** on the same thread
2. Select **Revoke Link**

The link becomes invalid immediately — giving you full control over external visibility and client-specific access.

***

#### **Delete Chat**

Deleting a chat permanently removes it from your workspace history. Use this to clear outdated drafts, test prompts, failed experiments, or irrelevant threads.

{% hint style="info" %}
Once deleted, the conversation cannot be recovered. Credits already consumed during that session are not refunded.
{% endhint %}


# Adding Context and Resources

How to improve chatbot responses by providing your own context through file uploads and knowledge bases.

### How Context Works in Qolaba

When you attach a file or a Knowledge Base to a chat, Qolaba processes and indexes that content using **Retrieval-Augmented Generation (RAG)** — a technique that allows the AI to retrieve relevant sections from your documents and use them as grounded context before generating a response.

This means the model isn't guessing or relying solely on training data. It is actively reading your material and anchoring its responses to what you've provided.

Once a file is uploaded to a chat, the AI retains context of it **throughout the entire session** — you don't need to keep the file selected or re-attach it with every prompt. Your uploaded files are also logged against the chat, so when you return to a previous conversation, you can always see which documents were used.

This is particularly valuable when:

* Working with proprietary information the model was never trained on
* Asking domain-specific questions that require your internal documents
* Maintaining consistency across tasks that rely on the same reference material
* Reducing the need to manually re-explain context in every prompt

***

### Two Ways to Add Context

1. [**Uploading Files**](/chatbot/adding-context-and-resources/uploading-files) — Attach files directly to a conversation for immediate, session-specific context. Best for one-off tasks where you need the AI to reference a specific document in that chat.
2. [**Knowledge Bases**](/chatbot/adding-context-and-resources/knowledge-bases) — Build persistent, reusable collections of documents powered by an upgraded RAG system for more accurate document retrieval and deeper context understanding. Best for recurring workflows where the same material is referenced repeatedly across chats, agents, or team members.

***

#### Which One Should You Use?

|                              | File Upload                      | Knowledge Base                           |
| ---------------------------- | -------------------------------- | ---------------------------------------- |
| **Best for**                 | One-off context, quick reference | Recurring workflows, team-shared context |
| **Persists across chats?**   | No — session only                | Yes — reusable across any chat           |
| **Supports multiple files?** | Yes                              | Yes                                      |
| **Shareable with team?**     | No                               | Yes                                      |
| **RAG-powered retrieval?**   | Yes                              | Yes — upgraded, more accurate            |

***

### Best Practices

* **One context per task** — Attach only what is relevant to the current prompt. Overloading the AI with unrelated documents reduces retrieval accuracy and response quality.
* **Name your files clearly** — The AI uses file names as part of its context. `brand-guidelines-2026.pdf` is more useful than `document1.pdf`.
* **Use Knowledge Bases for anything recurring** — If you are uploading the same file across multiple chats, it belongs in a Knowledge Base.
* **Trim large files before uploading** — Remove irrelevant pages or sections before uploading. Smaller, focused files improve retrieval precision and consume fewer credits.
* **Combine both methods when needed** — You can attach a Knowledge Base for persistent background context and upload a session-specific file for the task at hand. Both will be referenced together.
* **Review context before sending** — For complex tasks, confirm the right files or Knowledge Base are attached before running a prompt to avoid wasted credits.


# Uploading Files

How to upload files into a Qolaba chat, supported formats, and how to view and manage your upload history.

Upload files directly into a chat to give the LLM immediate context for that session. Qolaba processes your uploaded content using **RAG (Retrieval-Augmented Generation)**, allowing the model to retrieve and reference relevant sections from your documents throughout the conversation.

***

### Supported File Types

<table><thead><tr><th width="240.953125">File Type</th><th>Common Use</th></tr></thead><tbody><tr><td><strong>PDF</strong></td><td>Reports, research, documentation</td></tr><tr><td><strong>CSV</strong></td><td>Structured data, exports</td></tr><tr><td><strong>Excel (.xlsx)</strong></td><td>Spreadsheets, datasets</td></tr><tr><td><strong>DOCX</strong></td><td>Briefs, SOPs, brand guidelines</td></tr><tr><td><strong>TXT</strong></td><td>Notes, plain-text instructions</td></tr><tr><td><strong>Images</strong></td><td>Visual references, screenshots</td></tr></tbody></table>

***

### Uploading a File to Chat

**Step 1 —** Open a chat in your workspace.

**Step 2 —** Click the **attachment icon** or the **Plus (+) icon** in the prompt input area.

**Step 3 —** Select your file from your system or drag and drop it into the chat.

**Step 4 —** Wait for the upload to complete. The file name appears in the input area confirming it is attached.

**Step 5 —** Enter your prompt and send. The LLM will reference the uploaded file when generating its response.

{% hint style="info" %}
Once uploaded, the file remains active as context for the **entire chat session** — you do not need to keep it selected or re-attach it with each prompt.
{% endhint %}

***

### Viewing and Managing Uploaded Files

All files uploaded across your workspace — from chat sessions, Knowledge Bases, and other activities — are stored centrally in the **Files section** of the left navigation panel. This acts as a single repository for all your uploaded resources, making it easy to track, preview, and reuse files without uploading them again.

**To access:**

1. Go to the **left navigation panel**
2. Click the **File icon**

Inside, you can view the complete history of uploaded files including PDFs, DOCX, TXT, images, CSV, and other supported types.

***

#### **Previewing a File**

Click the **file name** to open a preview. Use this to verify content before reusing a file in a new chat or Knowledge Base.

***

#### **Deleting a File**

Click the **Delete option** next to the file to remove it from your upload history.

{% hint style="info" %}
If the file is currently used inside a Knowledge Base, deleting it from upload history will affect that Knowledge Base. Review dependencies before deleting.
{% endhint %}


# Knowledge Bases

What Knowledge Bases are, why to use them, supported formats, and how to access them in Qolaba.

A Knowledge Base is a structured, reusable collection of documents that the AI references before generating responses in chat. Instead of copying information into prompts repeatedly, upload your documents once — PDFs, DOCX, TXT files, images, or URLs — and attach the Knowledge Base to any chat or agent whenever that context is needed.

Qolaba's Knowledge Base system is powered by an upgraded **RAG (Retrieval-Augmented Generation)** engine, delivering more accurate document retrieval and deeper context understanding compared to standard file uploads.

***

### Why Use Knowledge Bases?

* **Context-aware responses** — The AI answers based on your actual documents, not just training data
* **No repetitive uploads** — Upload once, reuse across any chat or agent indefinitely
* **Domain-specific accuracy** — Ask questions directly about your reports, SOPs, brand guidelines, or research
* **Team consistency** — Share the same Knowledge Base across workspace members so everyone gets aligned outputs
* **Credit efficiency** — Avoid re-uploading the same files across multiple sessions

***

#### Example

A social media agency creates a Knowledge Base named `Competitor Research – Q1 2026` and uploads:

* Competitor case studies (PDF)
* Content strategy documents (DOCX)
* Analytics reports (PDF)
* Brand tone guidelines (TXT)

When a team member asks *"Create a content strategy based on competitor positioning"*, the model retrieves relevant sections from the uploaded material and grounds its response in that research — automatically, without any copy-pasting.

***

### Supported File Formats & Limits

<table><thead><tr><th width="281.60003662109375">File Type</th><th>Limit</th></tr></thead><tbody><tr><td><strong>PDF</strong></td><td>Up to 1,000 pages / 200 MB</td></tr><tr><td><strong>DOCX</strong></td><td>Up to 1,000 pages / 200 MB</td></tr><tr><td><strong>TXT</strong></td><td>Up to 200 MB</td></tr><tr><td><strong>Images</strong></td><td>Max 20 MB per image</td></tr><tr><td><strong>URLs</strong></td><td>Max 20 URLs per Knowledge Base</td></tr></tbody></table>

{% hint style="info" %}
Credits are consumed when uploading documents to a Knowledge Base. Credit deduction varies based on file size — larger files cost more to process.
{% endhint %}

***

#### Accessing Knowledge Bases

1. Go to the **left navigation panel**
2. Click the **Knowledge Base (File) icon**

This opens the Knowledge Base panel where you can view all existing Knowledge Bases or create a new one.


# Creating Knowledge Bases

Step-by-step guide to creating a Knowledge Base in Qolaba, uploading files, and selecting from your upload history.

#### Step 1 — Create a New Knowledge Base

1. Open the **Knowledge Base panel** from the left navigation panel
2. Click **Create New**
3. Enter a clear, descriptive name for your Knowledge Base — e.g., `Competitor Research – Q2 2026`
4. Click **Proceed**

***

#### Step 2 — Add Files

You will be prompted to add files immediately after creation. There are two ways to do this:

***

**Option A: Upload New Files**

Drag and drop files or click **Browse** to select from your system.

**Supported formats:** PDF, DOCX, TXT, Images

**Upload limits:**

| File Type | Limit                              |
| --------- | ---------------------------------- |
| Images    | Max 20 MB per image                |
| Documents | Max 1,000 pages or 200 MB per file |

> **Note:** Credits are consumed when uploading documents. Larger files cost more credits to process.

***

**Option B: Choose From Upload History**

Select previously uploaded files from your workspace without re-uploading them. Use the **Filter** option to narrow results by:

* Files
* URLs
* Images

**Selection limits:**

| Type           | Limit             |
| -------------- | ----------------- |
| Files and URLs | Max 20 selections |
| Images         | Max 10 selections |

***

#### Step 3 — Review and Confirm

Once files are selected, they appear at the top of the selection panel. Remove any incorrect selections using the **✕ icon**, then click **Add to Knowledge Base**.

Your Knowledge Base is now created and ready to use in any chat.


# Managing Knowledge Bases

How to view, edit, update, and delete Knowledge Bases and the files within them.

#### Viewing Your Knowledge Bases

1. Go to the **left navigation panel**
2. Click the **Knowledge Base (File) icon**

This opens a list of all your created Knowledge Bases. Select any one to manage it.

***

#### Available Actions

When you select a Knowledge Base, three actions are available:

* **Add to Chat** — Attach the Knowledge Base to your current chat session
* **Edit** — Enter the internal view to modify the Knowledge Base
* **Delete** — Permanently remove the Knowledge Base

***

#### Editing a Knowledge Base

Click the **Navigation icon** on a selected Knowledge Base to enter its internal view. From here you can:

* Rename the Knowledge Base
* Add more files
* Remove existing files
* Select specific files to attach to chat

***

**Adding More Files**

1. Click **Add Files**
2. Choose **Add New** to upload from your system, or **Choose From History** to select previously uploaded files
3. Remove any outdated files if needed
4. Save changes

**Example:** Your agency completed Q2 competitor research. Inside the existing Knowledge Base, remove the Q1 analytics reports, add the Q2 updated reports, and keep the brand strategy documents unchanged — without rebuilding the Knowledge Base from scratch.

***

**Removing Files**

1. Click on the file inside the Knowledge Base
2. Select the **remove/delete** option
3. Save changes

Keep your Knowledge Base lean and relevant — outdated files reduce retrieval accuracy.

***

#### Deleting a Knowledge Base

1. Select the Knowledge Base from the panel
2. Click **Delete**

This permanently removes the Knowledge Base. Files stored in your upload history are not affected and remain in your workspace unless deleted separately.


# Using Knowledge Bases

How to attach a Knowledge Base to a chat session and add context directly from the prompt input area.

#### Method 1 — From the Knowledge Base Panel

1. Open the **Knowledge Base panel** from the left navigation
2. Select the Knowledge Base you want to use
3. Click **Add to Chat**

A file icon appears in the chat interface confirming the Knowledge Base is active. The AI will now retrieve and reference your uploaded documents before generating every response in that session.

**Example:** A product team has a Knowledge Base named `Product Launch – Q3 2026` containing their PRD, competitive landscape report, ICP research, and positioning document. They attach it to chat and prompt: `Draft a go-to-market messaging framework based on our positioning and competitive gaps.`

Instead of producing a generic framework, the model retrieves the relevant sections from the PRD and competitive report and builds the response directly from the team's actual strategy — not assumptions.

***

#### Method 2 — From the Chat Input Area

You can also attach or create a Knowledge Base without leaving the chat interface:

1. Click the **Plus (+) icon** in the prompt input area
2. Select **Create Knowledge Base** to build a new one, or **Upload Knowledge Base** to attach an existing one

***

#### Attaching the Full Knowledge Base vs. Specific Files

When adding a Knowledge Base to chat, you can choose how much of it to attach:

1. **Attach the entire Knowledge Base** when your task requires broad context — strategic planning, long-form analysis, or research-heavy prompts where multiple documents are relevant. The model references all uploaded files.
2. **Attach specific files only** when your task is focused and doesn't need the full collection. Inside the Knowledge Base, select only the files relevant to your current prompt and attach those to chat.

**Example:** Your Knowledge Base contains:

* Brand tone guide
* Three competitor teardowns
* Customer research report
* Six months of campaign performance data

If you are writing a single email campaign, attach only:

* Brand tone guide
* Most recent campaign performance data

Attaching the full Knowledge Base when only two files are relevant adds unnecessary context, reduces retrieval precision, and may produce less accurate responses.


# Prompt and Controls

Every response the AI generates starts with a prompt. The quality, specificity, and structure of what you write directly determines the usefulness of what you get back. Beyond writing prompts, Qolaba gives you tools to save and reuse them — so recurring workflows don't require starting from scratch every time.

***

#### What's in This Section

1. [**Writing Your Prompt**](/chatbot/prompt-and-controls/writing-your-prompt) **→** Learn how to structure prompts for better outputs — covering the four elements of an effective prompt, best practices for iteration, and how to use the **Mic icon** for quick voice input when typing isn't the fastest option.
2. [**Saved Prompts**](/chatbot/prompt-and-controls/saved-prompts) **→** Store prompts you use regularly and insert them into any chat with one click. Covers creating, editing, copying, and deleting saved prompts — and when to use them for maximum efficiency.


# Writing your Prompt

How to write clear, effective prompts in Qolaba — including structure, best practices, and voice input.

#### Prompt Structure

A well-structured prompt typically includes four elements:

| Element           | What It Means                         |
| ----------------- | ------------------------------------- |
| **Objective**     | What you want the AI to do            |
| **Context**       | Background, audience, or purpose      |
| **Constraints**   | Tone, length, format, or limitations  |
| **Output format** | How the response should be structured |

**Example:**

Instead of:

```
Write a blog post.
```

Use:

```
Write a 600-word blog post about AI in retail.
Target audience: startup founders.
Tone: professional but conversational.
Include 3 subheadings and a short conclusion.
```

The second prompt gives the model everything it needs to produce a usable first draft — reducing back-and-forth and credit usage.

***

#### Best Practices

1. **Be specific** — Vague instructions produce vague outputs. Define exactly what you need.
2. **Provide context** — Mention the audience, purpose, background, or constraints relevant to the task.
3. **Define the output format** — Tell the model how to structure the response: bullet points, table, JSON, numbered list, paragraph, etc.
4. **Iterate rather than overload** — Start with a focused prompt and refine with follow-ups. Trying to cover everything in one prompt often produces unfocused responses.
5. **One objective per chat** — Start a new chat when switching to a completely different task. Mixing objectives in one thread affects context quality and output consistency.

***

#### Voice Input — Mic Icon

For quick or natural language input, use the **Mic icon** in the prompt input area. Click it and speak your prompt directly — Qolaba transcribes your speech into text, which you can review and edit before sending.

This is useful when:

* You want to describe a task naturally without typing out a structured prompt
* You are working quickly and want to capture an idea before refining it
* Typing a long, detailed prompt feels slower than speaking it


# Saved Prompts

How to create, manage, and use saved prompts in Qolaba for faster, more consistent workflows.

#### Accessing Saved Prompts

1. Go to the **left navigation panel**
2. Click the **Bookmark icon**

This opens your Saved Prompts library where you can view all existing prompts and create new ones.

***

#### Creating a Saved Prompt

1. Click **New Prompt**
2. Enter your prompt content
3. Save

The prompt appears in your library immediately and is ready to insert into any chat.

***

#### Managing Saved Prompts

**Editing a Prompt**

1. Locate the prompt in your library
2. Click **Edit**
3. Make your changes and click **Update**

**Deleting a Prompt**

1. Locate the prompt in your library
2. Click **Delete**

Deletion is permanent. The prompt is removed from your library immediately.

***

#### Using a Saved Prompt in Chat

**Insert into Chat**

1. Open **Saved Prompts** from the left navigation
2. Locate the prompt
3. Click **Add to Chat**

The prompt populates in your chat input field. Review and edit it if needed before sending.

**Copy a Prompt**

Click **Copy** on any saved prompt to copy it to your clipboard. Use this to reuse the prompt outside Qolaba or duplicate it as the basis for a new saved prompt.

***

#### What Saved Prompts Are Best For

* **Recurring task templates** — Weekly reports, content briefs, meeting summaries
* **Client-specific instructions** — Tone, format, and output requirements per client
* **Benchmarking prompts** — Consistent prompts used across model comparisons
* **Structured outputs** — Prompts that always require a specific format like JSON, tables, or outlines


# Model Selection

An overview of available model providers in Qolaba, their strengths and best use cases, and how to configure and compare models effectively.

Qolaba gives you access to multiple leading large language models from a single interface. Rather than being locked into one provider, you can choose the model that best fits your task — or switch between them to compare outputs.

***

#### Available Model Providers

Qolaba currently supports models from six providers. Each has distinct strengths — use this as a starting point when deciding which model to reach for:

| Provider               | Models                          | Best For                                                                     |
| ---------------------- | ------------------------------- | ---------------------------------------------------------------------------- |
| **GPT** (OpenAI)       | GPT-4o, GPT-4.1, o3, and others | General use, coding, structured outputs, reasoning                           |
| **Gemini** (Google)    | Gemini 2.0, 2.5 Pro and others  | Long context tasks, multimodal (text + vision), document analysis            |
| **Claude** (Anthropic) | Claude Sonnet, Opus and others  | Writing quality, nuanced reasoning, long-form content, instruction following |
| **DeepSeek**           | DeepSeek V3, R1 and others      | Technical reasoning, coding, cost-efficient performance                      |
| **Grok** (xAI)         | Grok 3, Grok 3 Mini and others  | Real-time information, conversational tasks, creative writing                |
| **Sonar** (Perplexity) | Sonar, Sonar Pro and others     | Web-grounded responses, research, fact-heavy queries                         |

{% hint style="info" %}
For full model specs, context lengths, capabilities, and plan availability, see [Model Reference→](/model-reference/chatbot-models) section.
{% endhint %}

***

#### Free vs. Paid Models

Model availability depends on your plan. Some models are accessible on all plans; others are available on paid plans only. See [Model Reference →](/model-reference/chatbot-models) for a complete breakdown.

***

#### What's in This Section

1. [**Model Information Panel →**](/chatbot/model-selection/model-information-panel) Understand what each model's information card tells you — context length, capability indicators, and credit costs.
2. [**Model Settings →**](/chatbot/model-selection/model-settings) Control how a model thinks and responds using Thinking Depth and Temperature.
3. [**Chat Branching →**](/chatbot/model-selection/chat-branching) Regenerate any response with a different model and compare outputs side-by-side within the same conversation.


# Model Information Panel

Understanding Model Information Panel in Qolaba — covering context length, capability indicators, and credit usage transparency.

#### 5.2.1 Context Length

* Token capacity
* Impact on performance

#### 5.2.2 Feature Indicators

* Text capability
* Vision capability
* Advanced reasoning capability

#### 5.2.3 Credit Usage Transparency

* Input credits per 1K tokens
* Output credits per 1K tokens
* How credits are calculated

***

## Model Information Breakdown

### 5.2.1 Context Length

This section explains how much information a model can handle in a single request.

#### • Token Capacity

* A **token** is a small unit of text (roughly 3–4 characters in English).
* Token capacity = maximum number of tokens the model can process at once.
* Includes:
  * User input
  * System instructions
  * Uploaded content
  * Model output

**Example:**&#x49;f a model supports 128K tokens, it can handle long documents, multi-step conversations, or large prompts in one go.

***

#### • Impact on Performance

Context length directly affects:

* **Memory** – Larger context = better understanding of long conversations.
* **Document handling** – Can process large PDFs, codebases, reports.
* **Cost** – Larger context models often cost more per request.
* **Latency** – Bigger context windows may slightly increase response time.

**In short:**&#x4D;ore context = more capability, but potentially higher cost and slower responses.

***

### 5.2.2 Feature Indicators

This section visually highlights what the model is capable of.

#### • Text Capability

Indicates the model can:

* Generate content
* Summarize
* Translate
* Code
* Answer questions

Almost all LLMs support text capability.

***

#### • Vision Capability

Indicates the model can:

* Analyze images
* Extract text from images
* Describe visuals
* Interpret charts or diagrams

If enabled, the model is multimodal (text + image).

***

#### • Advanced Reasoning Capability

Indicates stronger:

* Logical reasoning
* Multi-step problem solving
* Math and coding accuracy
* Complex decision-making

These models are typically:

* Slower than lightweight models
* More expensive
* Higher accuracy for complex tasks

***

### 5.2.3 Credit Usage Transparency

This section explains how model usage is billed.

#### • Input Credits per 1K Tokens

Cost charged for:

* Prompt text
* Uploaded files
* System instructions

Calculated per 1,000 tokens of input.

***

#### • Output Credits per 1K Tokens

Cost charged for:

* The model’s generated response

Also calculated per 1,000 tokens.

## Why This Structure Matters

* Compare models quickly
* Choose based on performance vs cost
* Understand technical limits before usage
* Avoid unexpected credit consumption


# Model Settings

How to use Thinking Depth and Temperature in Qolaba to control reasoning effort, thinking tokens, and response creativity.

Model Settings let you control how a model thinks and responds. Two settings are available — **Thinking Depth** and **Temperature** — each affecting a different dimension of the output.

***

#### 1. Thinking Depth

Thinking Depth controls how much internal reasoning the model applies before generating a response. When enabled, the model works through a series of reasoning steps — called **thinking tokens** — before producing its final answer.

**What Are Thinking Tokens?**

Thinking tokens are the model's internal reasoning steps — the process it goes through to interpret your prompt, evaluate different approaches, and arrive at a well-considered response before replying. They are visible in the response so you can follow how the model reasoned through your request.

Thinking tokens are counted as **output tokens** and consume credits accordingly. The deeper the thinking level, the more reasoning steps the model takes, and the more credits are used.

Thinking tokens are most valuable for:

* Complex research and analysis
* Multi-step problem solving
* Strategy planning and decision-making
* Advanced coding and debugging
* Tasks where understanding *how* the model reasoned matters as much as the answer itself

**Thinking Depth Levels**

| Level          | Reasoning Effort           | Credit Usage | Best For                                                               |
| -------------- | -------------------------- | ------------ | ---------------------------------------------------------------------- |
| **None**       | No thinking tokens         | Lowest       | Simple Q\&A, formatting, short rewrites                                |
| **Low**        | Minimal reasoning          | Low          | Basic content writing, casual prompts                                  |
| **Medium**     | Balanced reasoning         | Moderate     | Blog writing, coding assistance, structured business tasks             |
| **High**       | Deep, multi-step reasoning | High         | Complex coding, strategy planning, analytical writing                  |
| **Extra High** | Maximum reasoning effort   | Highest      | Advanced research, long-form reasoning chains, complex problem solving |

***

#### 2. Temperature

Temperature controls how creative or predictable the model's responses are — specifically, the randomness applied when the model selects words and constructs its response. In Qolaba, Temperature is set on a scale of **0 to 100**.

* **0** — Fully deterministic. The model picks the most probable word at every step. Responses are consistent, precise, and repeatable.
* **100** — Maximum randomness. The model explores less probable word choices, producing more varied, creative, and sometimes unexpected outputs.

<table><thead><tr><th width="105.53436279296875">Range</th><th width="283.85626220703125">Output Style</th><th>Best For</th></tr></thead><tbody><tr><td><strong>0 – 30</strong></td><td>Focused, deterministic, factual</td><td>Coding, legal drafts, data analysis, structured outputs (JSON, tables)</td></tr><tr><td><strong>40 – 60</strong></td><td>Balanced, natural, controlled</td><td>Blog posts, marketing copy, email drafts, general writing</td></tr><tr><td><strong>70 – 100</strong></td><td>Creative, varied, less predictable</td><td>Storytelling, brainstorming, ad copy, brand naming, ideation</td></tr></tbody></table>

***

#### Recommended Combinations

<table><thead><tr><th width="136.23748779296875">Thinking Depth</th><th width="132.08123779296875">Temperature</th><th>Output Type</th></tr></thead><tbody><tr><td><strong>High</strong></td><td><strong>0 – 30</strong></td><td>Precise and structured — technical reports, competitive analysis, complex coding</td></tr><tr><td><strong>Medium</strong></td><td><strong>40 – 60</strong></td><td>Balanced and reliable — business writing, content drafts, professional communication</td></tr><tr><td><strong>Low</strong></td><td><strong>70 – 100</strong></td><td>Fast and creative — brainstorming, ideation, headline generation, ad variations</td></tr></tbody></table>

**Examples:**

* Writing a detailed strategy document → **High Thinking Depth + Temperature 10–20**
* Generating 10 creative campaign name ideas → **Low Thinking Depth + Temperature 80–90**


# Chat Branching

### **What Is It?**

Chat Branching transforms the standard linear chat experience into a dynamic, multi-model conversation workspace. Users can regenerate any AI response using a different language model and continue the conversation from the reply they prefer.

This feature is a **model comparison and education tool**, enabling users to explore how different LLMs think, reason, and communicate side by side, all within the same chat interface.

{% embed url="<https://youtu.be/21-UtMCQ7MY?si=0_TwAA7FlGgfyOVt>" %}

***

### **Why It Matters**

Each language model has unique strengths—some excel at creative writing, others at analytical reasoning or concise summarization. Traditional chat interfaces lock users into one model per conversation, making it hard to compare responses.

Chat Branching removes this limitation. By allowing multiple regenerations of the same message, users can:&#x20;

* **Compare response quality** across models for the same prompt.
* **Identify model strengths** for specific tasks (e.g., coding, analysis, creative writing).
* **Build intuition** about model behavior, improving their ability to choose the right tool.
* **Explore alternative reasoning paths** without losing the original conversation flow.
* **Continue conversations** from any alternative response, pursuing the most promising line of thought.

This is a user-facing feature designed to help everyday users become more informed and effective AI consumers.


# How it Works ?

#### **Step 1: Have a Conversation**

Start a normal chat. Each assistant message is tagged with the model that generated it (e.g., "GPT-4.1", "Claude Sonnet 4.6"), so users always know which model they're interacting with.

#### **Step 2: Regenerate with a Different Model (Branching)**

Users can regenerate any assistant response using a different model. This creates a **branch**—an alternative response to the same question. The branch appears inline alongside the original, allowing users to compare them.

Users can generate up to **5 branches per message**, building a collection of responses from different models. For example:

<table data-header-hidden><thead><tr><th width="158.19921875">Branch</th><th width="265.453125">Model</th><th>Why</th></tr></thead><tbody><tr><td>Original</td><td>GPT-4.1</td><td>Default response</td></tr><tr><td>Branch 2</td><td>Claude Sonnet 4.6</td><td>Different writing style</td></tr><tr><td>Branch 3</td><td>Gemini 2.5 Pro</td><td>Compare Google's approach</td></tr><tr><td>Branch 4</td><td>DeepSeek R1</td><td>Reasoning-focused model</td></tr><tr><td>Branch 5</td><td>Grok 4</td><td>Another perspective</td></tr></tbody></table>

Each branch is clearly labeled with its model for easy comparison.

#### **Step 3: Continue from a Branch (Forking)**

After comparing branches, users can continue the conversation from any branch. This creates a **fork**—a new thread that picks up from the selected branch. The original thread remains intact, allowing users to revisit or explore other branches.

Forks can be nested up to **5 levels deep**, enabling exploratory conversation trees.

***

#### **Per-Message Branch Indicators**

Messages with branches display a count of alternative responses (e.g., "1 of 4"). Each response shows:&#x20;

* **Model name** (with logo).
* **Token usage** (response cost).
* **Reasoning content** (if supported by the model).

#### **Thread Sidebar**

Forked threads are **hidden from the main list** to keep the sidebar clean. They are accessible through the parent thread's conversation tree view.

#### **Full Conversation Tree**

The `/full` endpoint displays the entire conversation tree, showing where branches and forks occurred.




---

[Next Page](/llms-full.txt/1)

