# Blog
Source: https://docs.ai21.com/blog
# Changelog
Source: https://docs.ai21.com/changelog
## Export AI21 Documentation Pages as Markdown
**Release Date**: December 1, 2025
### What's New:
Documentation pages now include a dropdown button that lets you:
* Copy the current page as Markdown for easy use in AI tools or other workflows.
* View the page as plain-text Markdown directly.
This update makes it simpler for teams to reuse and share documentation content across tools and platforms.
## Together AI Open-Source Models Now Available in AI21 Maestro
**Release Date**: November 27, 2025
### What's New:
We’re excited to announce a new integration between **AI21 Labs** and **Together AI**, giving developers and enterprises **seamless access to Together AI’s open-source and frontier models directly within AI21 Maestro**.\
\
With this update, **AI21 Maestro now connects natively to Together AI’s AI-native cloud**, enabling teams to build knowledge agents with **greater flexibility, transparency, and cost efficiency**.
### Learn More:
For more details about this integration and what it enables, \
read the full announcement on our [blog](https://www.ai21.com/blog/ai21-together-ai-partnership/?utm_source=org-LI).
## Introducing Jamba Reasoning 3B
**Release Date**: October 8, 2025
### What's New:
We’re excited to announce Jamba Reasoning 3B, a compact open-source reasoning model that redefines what’s possible on-device.
Built on AI21 Labs’ novel SSM-Transformer hybrid architecture, this 3B-parameter model delivers **2–5× efficiency gains** over competitors while achieving **leading intelligence benchmarks**.
### Key Highlights
**Smarter, faster, smaller**: Hybrid SSM-Transformer efficiency enables long-context reasoning with minimal compute.
**Runs anywhere**: Deploy locally on your phone or computer via llama.cpp or LM Studio.
**Enterprise-ready**: Supports intelligent agentic workflows, reducing cloud costs and latency.
### Get Started:
* Check out the [AI21-Jamba-Reasoning-3B](https://huggingface.co/ai21labs/AI21-Jamba-Reasoning-3B) and [Jamba-Reasoning-3B-GGUF](https://huggingface.co/ai21labs/AI21-Jamba-Reasoning-3B-GGUF) cards.
* Read the full [blog post.](https://www.ai21.com/blog/introducing-jamba-reasoning-3B)
## You Can Now Use Requirements When Working With Tools!
**Release Date**: September 18, 2025
### What's New:
AI21 Maestro now supports Requirements with both File Search and Web Search. This allows you to define explicit constraints, such as format, tone, or content rules. AI21 Maestro will refine outputs to meet them, even when external tools are used for additional context.
### What This Means for You:
* More control over how responses are structured and styled
* Greater consistency and reliability, even when using external tools
* Streamlined workflows with outputs that meet your standards the first time
Check out the [API Reference](https://docs.ai21.com/reference/maestro-create-run#param-requirements).
## Deprecation Notice: tool\_resources Parameter
**Release Date**: September 18, 2025
### What's New:
The `tool_resources` parameter in POST /studio/v1/maestro/runs will be deprecated in **October 2025** . Please use the `tools` parameter instead.\
For more information, check the `tools` parameter[ documentation](https://docs.ai21.com/reference/maestro-create-run#param-tools).
## New Data Connectors Now Available for Enterprise Customers!
**Release Date**: September 16, 2025
### What's New:
* Available Data Connectors for Enterprise Users:
* Amazon S3
* Box Confluence
* Dropbox
* Google Drive
* Browse Website
* Intercom
* OneDrive
* Microsoft Teams
* Notion
* ServiceNow
* SharePoint
* Zendesk
Check out this [step-by-step guide](https://docs.ai21.com/docs/using-data-connectors#how-to-use-data-connectors%3A-step-by-step) for using Data Connectors in AI21 Studio.
## Mistral models are now available in AI21 Maestro!
**Release Date**: August 7, 2025
### What's New:
You can now access the `mistral-small`, `mistral-8x7b` and `mistral-7b` models directly in Maestro's API and playground. These models use AI21’s API key, so you don’t need to provide your own.\
\
See the [full list of available models](https://docs.ai21.com/reference/maestro-create-run#param-models).
## AI21 Maestro Is Now Live!
**Release Date**: July 8, 2025
We're excited to introduce AI21 Maestro.
### What's New:
AI21 Maestro is an advanced AI system for rapidly building and deploying knowledge agents that automate complex, high-value business tasks. Built for enterprise use, it combines retrieval-augmented generation (RAG), real-time validation, and dynamic execution planning to deliver accurate, reliable outputs aligned with user-defined goals and constraints.
**Validated Output** ensures that generated responses meet explicit user-defined requirements, such as formatting, content accuracy, tone, or domain-specific rules. Rather than relying solely on prompting, AI21 Maestro enforces these constraints by validating and automatically refining outputs using computational resources.
**AI21 Maestro RAG** features a robust RAG engine that enhances answer reliability by grounding model outputs in enterprise knowledge sources. It supports both semantic **file search** and **web search**, enabling contextual, fact-based responses from internal documents or real-time external content.
## Google Drive Connector Now Available for Enterprise Customers!
**Release Date**: July 8, 2025
You can now seamlessly import and sync your Google Drive files directly into your File Library, enabling faster, more integrated workflows.
### Key Benefits
* **Smart File Support** – Automatically imports commonly used formats: `.pdf`, `.docx`, `.txt`, `.html`, and `.md`.
* **Real-Time Sync** – Edits, renames, and newly added files in Google Drive are instantly reflected in the platform.
* **Setup Tip** – First-time connections may take longer for large accounts, as content is scanned and filtered during setup.
Our use of information received from Google Workspace APIs adheres to the [Google API Services User Data Policy](https://developers.google.com/terms/api-services-user-data-policy), including the Limited Use requirements. We do not use data obtained from these APIs to develop, improve, or train any machine learning or AI models beyond the user’s own personalized experience.
Want to learn more? [Talk to our sales team.](https://www.ai21.com/contact-sales/)
## Introducing Jamba 1.7!
**Release Date**: July 3, 2025
We're excited to introduce the release of Jamba 1.7!
### What's New:
* **Smarter Answers with Enhanced Grounding**: Jamba 1.7 now delivers more complete and accurate responses by better understanding context and focusing on what matters. Whether you're tackling complex questions or seeking precise insights, Jamba 1.7 is an ideal fit for question-answering and instruction-following tasks.
* **Faster and More Efficient**: Thanks to optimized configurations and quantization, Jamba 1.7 is delivering top-tier performance without sacrificing quality.
* **Self-Hosted Ready**: Want more control? Jamba 1.7 is now available for self-hosted deployment, giving teams flexibility and scalability in their own environments.\\
### Get Started:
Access via [AI21 Studio](https://studio.ai21.com/v2/workspaces/429d8ce9-0f35-4626-a48c-231d28db34a8/chat) or Hugging Face - [Jamba Large 1.7](https://huggingface.co/ai21labs/AI21-Jamba-Large-1.7), [Jamba Mini 1.7](https://huggingface.co/ai21labs/AI21-Jamba-Mini-1.7).
## Introducing Jamba 1.6!
**Release Date**: March 6, 2025
We're pleased to announce the release of Jamba 1.6.
### Key Updates:
* **Enhanced Model Quality**: Jamba Large 1.6 outperforms leading open models from Cohere, Meta, and Mistral on quality (Arena Hard) and speed.
* **Long context processing**: With a 256K context window and hybrid SSM-Transformer architecture, Jamba excels on efficiently and accurately processing long contexts, outperforming leading open model competitors on RAG and long context QA benchmarks
* **Secure Deployment**: Available via AI21 Studio (SaaS) or to download from Hugging Face and deploy privately (VPC/on-prem) from Hugging Face. More deployment options coming soon.
* **Improved Efficiency**: Faster response times with high accuracy.
### Get Started:
Access via [AI21 Studio](https://studio.ai21.com/v2/workspaces/429d8ce9-0f35-4626-a48c-231d28db34a8/chat) or Hugging Face - [Jamba Large 1.6](https://huggingface.co/ai21labs/AI21-Jamba-Large-1.6), [Jamba Mini 1.6](https://huggingface.co/ai21labs/AI21-Jamba-Mini-1.6).
# Community
Source: https://docs.ai21.com/community
# Cloud Platform Deployment
Source: https://docs.ai21.com/docs/cloud-platform-deployment
Deploy AI21's Jamba models on managed cloud services for production workloads
## Overview
Deploy AI21's Jamba models on managed cloud platforms for production-ready, scalable inference. Choose from the following cloud service options:
Deploy Jamba models using Amazon SageMaker's managed infrastructure.
Access Jamba models through Google Cloud's Vertex AI Model Garden.
For self-managed deployments, see our [Self Deployment Guide](/docs/self-deployment).
# Create an API key
Source: https://docs.ai21.com/docs/create-api-key
## Get an API key
This page explains how to create, view, and manage API keys for your workspace in [AI21 Studio](https://studio.ai21.com/v2?tab=maestro).
### **Step 1: Sign in to AI21 Studio**
You can sign in using your email address, Google account, GitHub account, or SSO.
### **Step 2: Open Settings**
In the bottom-left corner of the AI21 Studio sidebar, click **Settings**.
### Step 3: Open the API Keys tab
In the Workspace Settings panel, click the **API Keys** tab in the left-hand sidebar.
### Step 4: Create a New API Key
Click the **Create new key** button at the top-right corner of the Workspace Settings panel.
### Step 5: Name Your API Key
You may give your new key a name to help organize your workspace keys.\
If you prefer not to name it, leave the field blank and click **Create key**.\
To cancel, click **Cancel.**
### Step 6: Save Your API Key
Your new API key is now visible.\
Copy and save it in a secure location, **you won’t be able to view it again**.\
Anyone with your key can make requests on your behalf, so keep it private.\
If you lose it, you’ll need to generate a new one.
### Step 7: Manage Your API Keys
You can view and manage all workspace API keys in the **Workspace API Keys** section.\
You’ll see each key’s name and the last few characters (full keys are never shown again).\
From this page, you can rename or delete existing keys at any time.
# File Library
Source: https://docs.ai21.com/docs/file-library-studio
How to Use File Library in AI21 Studio
## Overview
The **File Library** in AI21 Studio serves as a centralized **knowledge base** for uploading, organizing, and managing organizational documents.
Once uploaded, these documents can be used to provide relevant context to Large Language Models (LLMs) when generating responses, enabling more accurate, grounded, and domain-aware outputs.
## Key Capabilities
* **Upload workspace documents** in supported formats (PDF, DOCX, HTML, TXT, Markdown)
* **Define data schemas** to enforce structure and consistency across similar document types
* **Search files** by name, date, ID, labels, or source
* **Select and filter documents** in the File Library
* **Edit file metadata**, including labels and paths
* **Delete outdated or unused files**
## **Upload a Document to File Library**
### **Step 1: Open the File Library Section**
Navigate to [File Library](https://studio.ai21.com/v2/workspaces/429d8ce9-0f35-4626-a48c-231d28db34a8/rag-engine) from the left-hand menu in AI21 Studio.
### **Step 2: Add a Document**
Click the **Add Document** button.
### **Step 3: Select One or More Files to Upload**
Supported file formats: **PDF, DOCX, TXT, HTML**, and **Markdown**.
You can upload up to **50 files,** each up to **100MB** in size, with a total limit of **1GB**.
### **Step 4: Label Your Documents**
Add new or existing labels to your documents, or leave the field blank to skip this step.\
Type labels and press enter to add.
### **Step 5: Connect**
Click **Connect** to start ingesting the data.
### **Step 6: View the New Files in File Library**
Once uploaded, the new documents will appear in the list.
##
## Search and Filter Files
### Step 1: Choose a Filter
Select how you want to filter the files from the dropdown menu. \
You can filter files by:
* Name
* Date
* ID
* Label
* Source
### Step 2: Search for Specific Documents
Fill in the search field based on your selected filter.\
\
For example:
* If you filter by **label**, the search field becomes a dropdown showing all available labels.
If you filter by **date**, enter a date, and the matching results will automatically appear.
### Optional: Select all documents
To select all documents in the File Library, click **Select All**.
## Delete Files From File Library
### **Step 1:** Locate and Select the File(s) You Want to Delete
In the table, select the file you want to delete from the \*\*Name \*\*column.
### **Step 2:** Delete the Selected Files
* On the right side of the table, click the (⋮) icon.
* Click **Delete**.
### **Step 3:** Confirm Deletion
A confirmation dialog will appear. Choose whether to delete the file or cancel the action.
### Step 4: View the Updated Document List
After deleting the selected file(s), the updated document list will be displayed in the File Library.
## **Next steps**
Check the API Refernce:
* [Upload Workspace Files](https://docs.ai21.com/reference/manage-library-ref/upload-workspace-files)
* [**List Library Files**](https://docs.ai21.com/reference/manage-library-ref/list-library-files)
* [**Update File**](https://docs.ai21.com/reference/manage-library-ref/update-file)
* [**Delete File**](https://docs.ai21.com/reference/manage-library-ref/delete-file)
Use the File Library in [AI21 Studio](https://studio.ai21.com/v2?tab=file_library).
# File Search
Source: https://docs.ai21.com/docs/file-search-studio-guide
How to Use File Search in AI21 Studio
Follow these steps to use documents to provide AI21 Maestro more context.
### Step 1: **Go to the Playground Section**
Navigate to Maestro [**Playground**](https://studio.ai21.com/v2?tab=maestro) from the left-hand menu in AI21 Studio.
### Step 2: Add a Tool
In the **Configuration** panel, go to the **Tools** section and click **Add**.
### **Step 3: Select File Search**
In the tools dropdown, select **File search**.
### **Step 4: Select Documents from Your File Library**
Choose the documents you want to use from your **File Library**.
When no files are selected, all files will be searched. Select docs to narrow your search.
### Save the Files You Want to Search
After verifying that the files are correct, click **Save**.
# Fine-tuning
Source: https://docs.ai21.com/docs/fine-tuning
Fine-tuning is the process of adapting a pre-trained model to perform better on specific tasks by training it on domain-specific data. Learn how to fine-tune Jamba models using different approaches including full fine-tuning, LoRA, and QLoRA
## Overview
The Jamba models can be fine-tuned using several approaches:
* **Full Fine-tuning**: Complete model parameter updates (requires significant GPU resources)
* **LoRA (Low-Rank Adaptation)**: Parameter-efficient fine-tuning approach
* **QLoRA**: Combines LoRA with 4-bit quantization for single GPU training
## Full Fine-tuning
Full fine-tuning updates all model parameters and provides the most comprehensive training results.
For a comprehensive implementation guide using AWS SageMaker with multi-node and FSDP configuration, see the [AI21 SageMaker Fine-tuning Repository](https://github.com/AI21Labs/hf-finetune-sagemaker).
Full fine-tuning requires multiple high-memory GPUs.
## LoRA Fine-tuning
LoRA (Low-Rank Adaptation) fine-tuning injects compact, low-rank adapter layers into a frozen pretrained model—letting you specialize it for your task with just a few percent of the parameters, minimal extra compute and storage and with a small loss in accuracy or inference speed.
### Prerequisites
Before starting LoRA fine-tuning, install the required dependencies:
```bash theme={"system"}
pip install trl transformers torch datasets peft
```
This LoRA fine-tuning example uses bfloat16 precision and requires \~130GB GPU RAM (e.g., 2x A100 80GB GPUs).
### Implementation
```python theme={"system"}
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig
from peft import LoraConfig
tokenizer = AutoTokenizer.from_pretrained("ai21labs/AI21-Jamba-Mini-1.7")
model = AutoModelForCausalLM.from_pretrained(
"ai21labs/AI21-Jamba-Mini-1.7",
device_map="auto", # Automatically distribute across available GPUs
torch_dtype=torch.bfloat16, # Use mixed precision for memory efficiency
attn_implementation="flash_attention_2", # Optimized attention implementation
)
```
```python theme={"system"}
lora_config = LoraConfig(
r=8, # Rank of adaptation - controls the number of trainable parameters
target_modules=[
"embed_tokens",
"x_proj", "in_proj", "out_proj", # mamba layers
"gate_proj", "up_proj", "down_proj", # mlp layers
"q_proj", "k_proj", "v_proj", "o_proj", # attention layers
],
task_type="CAUSAL_LM",
bias="none",
)
```
```python theme={"system"}
# Load dataset (replace with your own dataset)
dataset = load_dataset("philschmid/dolly-15k-oai-style", split="train")
```
```python theme={"system"}
training_args = SFTConfig(
output_dir="/dev/shm/results", # Where to save the model
logging_dir="./logs", # Where to save training logs
num_train_epochs=2, # Number of training epochs
per_device_train_batch_size=4, # Batch size per GPU
learning_rate=1e-5, # Learning rate for fine-tuning
logging_steps=10, # Log training metrics every 10 steps
gradient_checkpointing=True, # Save memory at cost of compute
max_seq_length=4096, # Maximum sequence length
save_steps=100, # Save model checkpoint every 100 steps
)
```
```python theme={"system"}
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
args=training_args,
peft_config=lora_config,
train_dataset=dataset,
)
trainer.train()
```
The dataset in this example uses conversational format (with `messages` column), so `SFTTrainer` automatically applies Jamba's chat template. For more information about supported dataset formats and advanced SFTTrainer features, see the [TRL documentation](https://huggingface.co/docs/trl/main/en/sft_trainer#dataset-format-support).
For Jamba Large LoRA fine-tuning, we recommend using the qLoRA+FSDP approach detailed in the QLoRA section below, as it provides better memory efficiency for the larger model.
## QLoRA Fine-tuning
[QLoRA](https://arxiv.org/abs/2305.14314) combines LoRA with 4-bit quantization, making it possible to fine-tune on a single 80GB GPU while maintaining good performance.
### Prerequisites
Before starting QLoRA fine-tuning, install the required dependencies:
```bash theme={"system"}
pip install trl transformers torch datasets peft bitsandbytes
```
### Implementation
```python theme={"system"}
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig
from peft import LoraConfig
tokenizer = AutoTokenizer.from_pretrained("ai21labs/AI21-Jamba-Mini-1.7")
# Configure 4-bit quantization
quantization_config = BitsAndBytesConfig(
load_in_4bit=True, # Enable 4-bit quantization
bnb_4bit_quant_type="nf4", # Use NormalFloat 4-bit quantization
bnb_4bit_compute_dtype=torch.bfloat16, # Compute in bfloat16 for better stability
)
```
```python theme={"system"}
model = AutoModelForCausalLM.from_pretrained(
"ai21labs/AI21-Jamba-Mini-1.7",
device_map="auto", # Automatically distribute across available GPUs
quantization_config=quantization_config, # Apply 4-bit quantization
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
)
```
```python theme={"system"}
lora_config = LoraConfig(
r=8, # Rank of adaptation - controls trainable parameters
target_modules=[
"embed_tokens",
"x_proj", "in_proj", "out_proj", # mamba layers
"gate_proj", "up_proj", "down_proj", # mlp layers
"q_proj", "k_proj", "v_proj", "o_proj", # attention layers
],
task_type="CAUSAL_LM",
bias="none",
)
```
```python theme={"system"}
# Load dataset (replace with your own dataset)
dataset = load_dataset("philschmid/dolly-15k-oai-style", split="train")
```
```python theme={"system"}
training_args = SFTConfig(
output_dir="./results", # Where to save the model
logging_dir="./logs", # Where to save training logs
num_train_epochs=2, # Number of training epochs
per_device_train_batch_size=8, # Higher batch size possible with quantization
learning_rate=1e-5, # Learning rate for fine-tuning
logging_steps=1, # Log training metrics every step
gradient_checkpointing=True, # Save memory at cost of compute
gradient_checkpointing_kwargs={"use_reentrant": False}, # Required for some models
save_steps=100, # Save model checkpoint every 100 steps
max_seq_length=4096, # Maximum sequence length
)
```
```python theme={"system"}
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
args=training_args,
peft_config=lora_config,
train_dataset=dataset,
)
trainer.train()
```
Jamba Large fine-tuning requires 8x A100/H100 80GB GPUs and uses qLoRA+FSDP. This approach uses axolotl framework with a modified transformers version to optimize memory usage.
Due to its size, in order to run the training on a single 8 GPU node, Jamba Large 1.7 has to be quantized. This can happen either at the start of the training job, or in a pre-process step. If you want to pre-quantize the model, you can do that easily using bitsandbytes (make sure to use `bnb_4bit_quant_storage=torch.bfloat16` so you can use FSDP).
```bash theme={"system"}
# Install axolotl and dependencies
git clone https://github.com/axolotl-ai-cloud/axolotl
cd axolotl
pip3 install packaging ninja
pip3 install -e '.[flash-attn,deepspeed]'
pip install bitsandbytes~=0.43.3
pip install trl
pip install peft~=0.12.0
pip install accelerate~=0.33.0
pip install mamba-ssm causal-conv1d>=1.2.0
pip install git+https://github.com/xgal/transformers@897f80665c37c531b7803f92655dbc9b3a593fe7
```
```python theme={"system"}
from transformers import AutoConfig, AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch
model_name = "ai21labs/AI21-Jamba-Large-1.7"
quantization_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_compute_dtype=torch.bfloat16,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_quant_storage=torch.bfloat16,
)
quantized_model = AutoModelForCausalLM.from_pretrained(
model_name,
device_map="auto",
torch_dtype=torch.bfloat16,
quantization_config=quantization_config
)
tokenizer = AutoTokenizer.from_pretrained(model_name)
tokenizer.save_pretrained('AI21-Jamba-Large-1.7-BNB-nf4-bf16')
quantized_model.save_pretrained('AI21-Jamba-Large-1.7-BNB-nf4-bf16')
```
```bash theme={"system"}
# Run training with axolotl (change base_model to pre-quantized model if using pre-quantization)
accelerate launch -m axolotl.cli.train examples/jamba/qlora_fsdp.yaml
```
For detailed configuration files and examples, visit the [axolotl Jamba examples](https://github.com/axolotl-ai-cloud/axolotl/tree/main/examples/jamba). The modified transformers version prevents excessive CPU RAM usage that would otherwise require over 1.6TB instead of the required 200GB.
# Function Calling
Source: https://docs.ai21.com/docs/function-calling
Learn how to extend Jamba's capabilities by giving it access to custom functions and external tools
## Overview
Function calling allows Jamba models to intelligently decide when and how to call external functions or tools based on the conversation context. Instead of trying to answer everything directly, the model can request specific function calls with appropriate parameters, enabling integration with APIs, databases, calculators, and other external systems.
Function calling is available for all Jamba models via the Chat Completions API.
## How It Works
1. **Define your functions** - Provide JSON schemas describing available tools
2. **Jamba decides** - The model determines if and when to call functions
3. **Execute functions** - Your code runs the requested functions
4. **Get response** - Jamba uses the results to generate its final answer
## Function Calling Example
First, define the function you want to make available to the model:
```python Python theme={"system"}
import json
import requests
def convert_currency(amount: float, from_currency: str, to_currency: str) -> str:
"""Convert an amount from one currency to another using live exchange rates"""
# Using exchangerate-api.com (free tier available)
response = requests.get(
f"https://api.exchangerate-api.com/v4/latest/{from_currency.upper()}"
)
data = response.json()
if to_currency.upper() not in data["rates"]:
return f"Currency {to_currency} not found"
exchange_rate = data["rates"][to_currency.upper()]
converted_amount = amount * exchange_rate
result = {
"original_amount": amount,
"from_currency": from_currency.upper(),
"to_currency": to_currency.upper(),
"exchange_rate": exchange_rate,
"converted_amount": round(converted_amount, 2)
}
return json.dumps(result)
# Define the function schema for the model
tools = [
{
"type": "function",
"function": {
"name": "convert_currency",
"description": "Convert an amount from one currency to another using current exchange rates",
"parameters": {
"type": "object",
"properties": {
"amount": {
"type": "number",
"description": "The amount to convert"
},
"from_currency": {
"type": "string",
"description": "The source currency code (e.g., USD, EUR, GBP)"
},
"to_currency": {
"type": "string",
"description": "The target currency code (e.g., USD, EUR, GBP)"
}
},
"required": ["amount", "from_currency", "to_currency"]
}
}
}
]
```
```typescript TypeScript theme={"system"}
async function convertCurrency(amount: number, fromCurrency: string, toCurrency: string): Promise {
// Convert an amount from one currency to another using live exchange rates
const response = await fetch(
`https://api.exchangerate-api.com/v4/latest/${fromCurrency.toUpperCase()}`
);
const data = await response.json();
if (!(toCurrency.toUpperCase() in data.rates)) {
return `Currency ${toCurrency} not found`;
}
const exchangeRate = data.rates[toCurrency.toUpperCase()];
const convertedAmount = amount * exchangeRate;
const result = {
original_amount: amount,
from_currency: fromCurrency.toUpperCase(),
to_currency: toCurrency.toUpperCase(),
exchange_rate: exchangeRate,
converted_amount: Math.round(convertedAmount * 100) / 100
};
return JSON.stringify(result);
}
// Define the function schema for the model
const tools = [
{
type: "function",
function: {
name: "convert_currency",
description: "Convert an amount from one currency to another using current exchange rates",
parameters: {
type: "object",
properties: {
amount: {
type: "number",
description: "The amount to convert"
},
from_currency: {
type: "string",
description: "The source currency code (e.g., USD, EUR, GBP)"
},
to_currency: {
type: "string",
description: "The target currency code (e.g., USD, EUR, GBP)"
}
},
required: ["amount", "from_currency", "to_currency"]
}
}
}
];
```
Create a chat request with your tools and user message:
```python Python theme={"system"}
from ai21 import AI21Client
from ai21.models.chat import ChatMessage
client = AI21Client()
messages = [
ChatMessage(
role="user",
content="How much is 100 USD in euros?"
)
]
response = client.chat.completions.create(
model="jamba-large",
messages=messages,
tools=tools,
max_tokens=150
)
print(response.choices[0].message)
```
```typescript TypeScript theme={"system"}
import { AI21Client } from 'ai21-typescript';
const client = new AI21Client();
const messages = [
{
role: "user",
content: "Convert 500 GBP to Japanese yen"
}
];
const response = await client.chat.completions.create({
model: "jamba-large",
messages: messages,
tools: tools,
max_tokens: 150
});
console.log(response.choices[0].message);
```
Check if the model wants to call a function and execute it:
```python Python theme={"system"}
# Check if the model wants to call a function
message = response.choices[0].message
if message.tool_calls:
# Execute the requested function
for tool_call in message.tool_calls:
function_name = tool_call.function.name
function_args = json.loads(tool_call.function.arguments)
if function_name == "convert_currency":
function_result = convert_currency(**function_args)
# Add the function call and result to the conversation
messages.append(message) # Assistant's function call
messages.append(ChatMessage(
role="tool",
tool_call_id=tool_call.id,
content=function_result
))
```
```typescript TypeScript theme={"system"}
// Check if the model wants to call a function
const message = response.choices[0].message;
if (message.tool_calls) {
// Execute the requested function
for (const toolCall of message.tool_calls) {
const functionName = toolCall.function.name;
const functionArgs = JSON.parse(toolCall.function.arguments);
if (functionName === "convert_currency") {
const functionResult = await convertCurrency(
functionArgs.amount,
functionArgs.from_currency,
functionArgs.to_currency
);
// Add the function call and result to the conversation
messages.push(message); // Assistant's function call
messages.push({
role: "tool",
tool_call_id: toolCall.id,
content: functionResult
});
}
}
}
```
Send the function results back to the model for the final response:
```python Python theme={"system"}
# Get the final response with function results
final_response = client.chat.completions.create(
model="jamba-large",
messages=messages,
tools=tools,
max_tokens=150
)
print(final_response.choices[0].message.content)
else:
# No function call needed, model responded directly
print(message.content)
```
```typescript TypeScript theme={"system"}
// Get the final response with function results
const finalResponse = await client.chat.completions.create({
model: "jamba-large",
messages: messages,
tools: tools,
max_tokens: 150
});
console.log(finalResponse.choices[0].message.content);
} else {
// No function call needed, model responded directly
console.log(message.content);
}
```
## Complete Examples
For reference, here are the complete working examples that put all the steps together:
```python theme={"system"}
import json
import requests
from ai21 import AI21Client
from ai21.models.chat import ChatMessage
def convert_currency(amount: float, from_currency: str, to_currency: str) -> str:
"""Convert an amount from one currency to another using live exchange rates"""
# Using exchangerate-api.com (free tier available)
response = requests.get(
f"https://api.exchangerate-api.com/v4/latest/{from_currency.upper()}"
)
data = response.json()
if to_currency.upper() not in data["rates"]:
return f"Currency {to_currency} not found"
exchange_rate = data["rates"][to_currency.upper()]
converted_amount = amount * exchange_rate
result = {
"original_amount": amount,
"from_currency": from_currency.upper(),
"to_currency": to_currency.upper(),
"exchange_rate": exchange_rate,
"converted_amount": round(converted_amount, 2)
}
return json.dumps(result)
# Define the function schema for the model
tools = [
{
"type": "function",
"function": {
"name": "convert_currency",
"description": "Convert an amount from one currency to another using current exchange rates",
"parameters": {
"type": "object",
"properties": {
"amount": {
"type": "number",
"description": "The amount to convert"
},
"from_currency": {
"type": "string",
"description": "The source currency code (e.g., USD, EUR, GBP)"
},
"to_currency": {
"type": "string",
"description": "The target currency code (e.g., USD, EUR, GBP)"
}
},
"required": ["amount", "from_currency", "to_currency"]
}
}
}
]
# Initialize client and create initial message
client = AI21Client()
messages = [
ChatMessage(
role="user",
content="How much is 100 USD in euros?"
)
]
# Make initial request
response = client.chat.completions.create(
model="jamba-large",
messages=messages,
tools=tools,
max_tokens=150
)
message = response.choices[0].message
# Check if the model wants to call a function
if message.tool_calls:
# Execute the requested function
for tool_call in message.tool_calls:
function_name = tool_call.function.name
function_args = json.loads(tool_call.function.arguments)
if function_name == "convert_currency":
function_result = convert_currency(**function_args)
# Add the function call and result to the conversation
messages.append(message) # Assistant's function call
messages.append(ChatMessage(
role="tool",
tool_call_id=tool_call.id,
content=function_result
))
# Get the final response with function results
final_response = client.chat.completions.create(
model="jamba-large",
messages=messages,
tools=tools,
max_tokens=150
)
print(final_response.choices[0].message.content)
else:
# No function call needed, model responded directly
print(message.content)
```
```typescript theme={"system"}
import { AI21Client } from 'ai21-typescript';
async function convertCurrency(amount: number, fromCurrency: string, toCurrency: string): Promise {
// Convert an amount from one currency to another using live exchange rates
const response = await fetch(
`https://api.exchangerate-api.com/v4/latest/${fromCurrency.toUpperCase()}`
);
const data = await response.json();
if (!(toCurrency.toUpperCase() in data.rates)) {
return `Currency ${toCurrency} not found`;
}
const exchangeRate = data.rates[toCurrency.toUpperCase()];
const convertedAmount = amount * exchangeRate;
const result = {
original_amount: amount,
from_currency: fromCurrency.toUpperCase(),
to_currency: toCurrency.toUpperCase(),
exchange_rate: exchangeRate,
converted_amount: Math.round(convertedAmount * 100) / 100
};
return JSON.stringify(result);
}
// Define the function schema for the model
const tools = [
{
type: "function",
function: {
name: "convert_currency",
description: "Convert an amount from one currency to another using current exchange rates",
parameters: {
type: "object",
properties: {
amount: {
type: "number",
description: "The amount to convert"
},
from_currency: {
type: "string",
description: "The source currency code (e.g., USD, EUR, GBP)"
},
to_currency: {
type: "string",
description: "The target currency code (e.g., USD, EUR, GBP)"
}
},
required: ["amount", "from_currency", "to_currency"]
}
}
}
];
async function functionCallingExample() {
// Initialize client and create initial message
const client = new AI21Client();
const messages = [
{
role: "user",
content: "Convert 500 GBP to Japanese yen"
}
];
// Make initial request
const response = await client.chat.completions.create({
model: "jamba-large",
messages: messages,
tools: tools,
max_tokens: 150
});
const message = response.choices[0].message;
// Check if the model wants to call a function
if (message.tool_calls) {
// Execute the requested function
for (const toolCall of message.tool_calls) {
const functionName = toolCall.function.name;
const functionArgs = JSON.parse(toolCall.function.arguments);
if (functionName === "convert_currency") {
const functionResult = await convertCurrency(
functionArgs.amount,
functionArgs.from_currency,
functionArgs.to_currency
);
// Add the function call and result to the conversation
messages.push(message); // Assistant's function call
messages.push({
role: "tool",
tool_call_id: toolCall.id,
content: functionResult
});
}
}
// Get the final response with function results
const finalResponse = await client.chat.completions.create({
model: "jamba-large",
messages: messages,
tools: tools,
max_tokens: 150
});
console.log(finalResponse.choices[0].message.content);
} else {
// No function call needed, model responded directly
console.log(message.content);
}
}
// Run the example
functionCallingExample().catch(console.error);
```
# HTTP Tools
Source: https://docs.ai21.com/docs/http-studio-guide
How to Use HTTP Tools in AI21 Studio
HTTP tools allow AI21 Maestro to call external services directly using standard HTTP requests. They’re best for simple or one-off integrations that don’t require a full MCP server.\
Follow these steps to give your agent access to external APIs.
### Step 1: **Go to the Playground Section**
Navigate to Maestro [**Playground**](https://studio.ai21.com/v2?tab=maestro) from the left-hand menu in AI21 Studio.
### Step 2: Add a Tool
In the **Configuration** panel, go to the **Tools** section and click **Add**.
### **Step 3: Select HTTP Tools**
In the tools dropdown, select **MCP**.
### **Step 4: Fill in the Fields**
**4.1 Define the JSON Schema**\
Provide the **JSON schema** that describes the API operation you want to connect.
* This schema defines the function name, description, parameters, and their types.
* Use the exact schema of the API you want your agent to call.
**4.2 Enter the URL**\
Add the **API endpoint URL** that the tool should call.\
Use the full URL, for example: `https://api.example.com/weather`.
**4.3 Add the Authorization Header** \
If your API requires authentication, enter the **authorization header** \
(for example, `Bearer {api_key}`).
AI21 Maestro does not store your API credentials or any provided header.
### **Step 5: Add the HTTP Tool**
After verifying the tools are correct, click **Add**.
### **Step 6 (Optional): Edit HTTP Tools**
Edit an existing MCP tool by following these steps:
1. Click the **⋮** (colon icon).
2. Select **Edit**, make your changes, and click **Save.**
### Next Steps
Use HTTP tools in [AI21 Stuido](https://studio.ai21.com/v2?tab=maestro).
# Quick Start
Source: https://docs.ai21.com/docs/instruction-following-module
This guide focuses on AI21 Maestro's Validated Output.
## What is AI21 Maestro’s Validated Output?
One of AI21 Maestro's core capabilities is providing validated output, which addresses a critical problem that even advanced language models can struggle with - consistently following complex instructions that include multiple constraints.
AI21 Maestro ensures your language model's outputs meet your specific requirements. Instead of relying on prompts to work reliably on their own, it uses computational resources to validate and refine outputs until they satisfy your constraints.
AI21 Maestro's Validated Output addresses this by:
* **Validating** outputs against your explicit requirements
* **Automatically fixes outputs** that don't meet the requirements
* **Provides a report** on requirement fulfillment with detailed scores
## Key Concepts
* **Requirements**: Explicit constraints you define for your outputs (e.g., format, tone, content rules).
* **Budget**: Controls computational effort and trade-offs—higher budgets use more resources (increasing cost and latency) to achieve better adherence to requirements (high/medium/low).
* **Requirements Report**: Provides detailed scoring and feedback on how well each requirement was met.
* **Model Agnostic**: Maestro works with both AI21's first-party models and popular third-party models (e.g., GPT-4, Claude, Gemini)—choose the model that best fits your needs.
## Basic Usage
**Simple Example**
```python theme={"system"}
import os
from ai21 import AI21Client
client = AI21Client(api_key=os.getenv("AI21_API_KEY"))
# create_and_poll() returns after the processing ended or the default timeout is passed
run = client.beta.maestro.runs.create_and_poll(
input="Write a Python function to calculate fibonacci numbers",
requirements=[
{
"name": "function_length",
"description": "The function should be no more than 10 lines long"
},
{
"name": "include_docstring",
"description": "Include a Google-style docstring explaining the function"
}
],
budget="low",
include=["requirements_result"]
)
# Result is available immediately when this returns
print(f"Result: {run.result}")
print(f"Requirements Score: {run.requirements_result.score}")
```
**Understanding Asynchronous Execution**
Maestro runs execute asynchronously in the backend. The **create()** method returns immediately with a run ID, while processing happens in the background:
```python theme={"system"}
import os
import time
from ai21 import AI21Client
client = AI21Client(api_key=os.getenv("AI21_API_KEY"))
# create() returns immediately, processing happens asynchronously
run = client.beta.maestro.runs.create(
input="Write a marketing email",
requirements=[{"name": "word_count", "description": "use 150-200 words"}],
budget="low"
)
print(f"Run ID: {run.id}") # Available immediately
print(f"Status: {run.status}") # "in_progress"
# Poll for completion manually
while run.status == "in_progress":
time.sleep(5)
run = client.beta.maestro.runs.retrieve(run.id)
print(f"Final result: {run.result}")
```
## Budget Levels
* **Low**: Lower latency and lighter approach for getting a validated output
* **Medium**: Balanced approach with moderate fix attempts
* **High**: Maximum reliability with multiple fix attempts and parallel processing
## Quick Tips
1. **Be Specific**: Write clear, measurable requirements
2. **Start Small**: Begin with 2-3 requirements and expand
3. **Use the Playground**: Test your requirements in our web interface before implementing
4. **Check Scores**: Review requirement scores to understand what's working
## Next Steps
* Log in to [AI21 Studio](https://studio.ai21.com/v2?tab=maestro) and experiment the Maestro Playground.
* Set up a 3rd party model in AI21 Studio's[ Model Integrations Page](https://studio.ai21.com/v2?tab=third_party_models).
* Read the [Complete Walkthrough Guide](https://docs.ai21.com/docs/walkthrough-guide) for advanced usage patterns.
* Explore the [API Reference](https://docs.ai21.com/reference/endpoints) for full parameter details.
* If you have any technical questions about Maestro, feel free to reach out to our support team via [email](mailto:support@ai21.com) or click the chat icon in the lower right corner.
# Batch API
Source: https://docs.ai21.com/docs/jamba-batch-api
## Overview
Accelerate high-volume AI workflows with AI21's Batch API, designed to support enterprise-scale use cases where traditional real-time APIs fall short.
Many enterprises need to process tens of thousands of language model requests quickly and cost-effectively. Use cases like product data enrichment, large-scale content classification, and historical knowledge audits demand a solution that’s purpose-built for asynchronous, high-throughput jobs.
AI21’s Batch API helps deliver rapid results at scale without the need for complex infrastructure or custom scripting.
## How it Works
Instead of sending each request individually through a synchronous API, the Batch API allows you to submit a `.jsonl` file containing thousands of inputs in a single job. The requests are processed asynchronously, and the results are returned in a downloadable output file.
### Workflow
1. Prepare your `.jsonl` input file
Each line in the file represents a single prompt or input payload.
2. Submit the batch request
Use a simple HTTP request to submit the file along with your desired endpoint.
3. Monitor and manage jobs
Query the status and progress of your batch jobs and retrieve results. The result file is available once the batch is finished.
4. Download results
When the batch completes, download your results directly.
## Key Advantages of AI21’s Batch API
* **Enterprise-scale throughput**
Submit and process thousands of requests in a single job.
* **Simple integration**
Connect easily to your existing AI workflows via AI21’s SaaS infrastructure.
* **Battle-tested in production**
Built in partnership with our customers, where it has cut processing time for large classification jobs from several hours to under one hour.
## Get Started
AI21’s Batch API is currently available for enterprise use through our SaaS offering. To request access or explore how Batch can integrate into your workflows, please [contact our sales team](https://www.ai21.com/contact-sales/). We’ll help tailor a solution that fits your business goals.
# Jamba
Source: https://docs.ai21.com/docs/jamba-foundation-models
Introducing the Jamba Family of Open Models.
## Overview
Built on the novel Mamba-Transformer architecture, these highly efficient and powerful models push the boundaries of AI. They deliver unmatched speed and quality and feature the longest context window (256K tokens) among open models.
Our most powerful and advanced model, designed to handle complex tasks at enterprise scale with superior performance.
Jamba2 Mini blends efficiency and steerability into a 12B-active parameters model, delivering reliable output on core enterprise workflows.
Jamba2 3B packs reliability and steerability into a compact 3B model, powering on-device applications and supporting agentic systems.
## Key Benefits
Retain total control over your data, with zero data visibility for the model vendor by deploying Jamba privately in your VPC or on-premises. Ideal for organizations or use cases handling regulated data (e.g., in finance or healthcare) or confidential proprietary data. With comparable quality to market-leading models, you get superior quality with total data privacy.
With the weights available to be downloaded into your environment, Jamba can be customized to your domain and use case. This is a model that's fully yours.
With a 256K context window, Jamba excels on the kinds of use cases enterprises rely on in their workflows, from analyzing lengthy documents to enhancing RAG workflows at the retrieval stage. And due to its efficiency-optimized hybrid architecture, it can do all this at a lower cost than competitors.
## Self Deployment Options
Jamba models are ideal for enterprises that need to maintain full control over data and performance. Deployment options include:
* **Cloud-hosted VPC** – Isolate and scale your deployment in a virtual private cloud environment.
* **On-premises environments** – Run Jamba entirely within your own infrastructure to meet strict compliance or latency requirements.
* **Custom hybrid solutions** – Combine cloud and on-premises capabilities to fit your architecture, performance, and governance needs.
## Supported Languages
Jamba models officially support 9 languages:
* English, Spanish, French, Portuguese, Italian, Dutch, German, Arabic, Hebrew.
## Model Details
| Model | Model Size | Max Tokens | Version | Snapshot | API Endpoint |
| ----------- | ---------------------------------- | ---------- | ------- | -------- | ------------- |
| Jamba Large | 398B parameters (94B active) | 256K | 1.7 | 2025-07 | `jamba-large` |
| Jamba Mini | 52B parameters (12B active) | 256K | 2 | 2026-01 | `jamba-mini` |
| Jamba 3B | 3B parameters | 256K | 2 | 2026-01 | N/A |
Engineers and data scientists at AI21 labs created the model to help developers and businesses leverage AI to build real-world products with tangible value. The Jamba model family supports multi-language and zero-shot instruction-following.
* **Organization developing model:** AI21 Labs
* **Model type:** Joint Attention and Mamba (Jamba)
* **Knowledge cutoff date** August 22nd, 2024
* **Input Modality:** Text
* **Output Modality:** Text
* **Contact:** [info@ai21.com](mailto:info@ai21.com)
## API Versioning
We advise using dated versions of the Jamba API to avoid disruptions from model updates and breaking changes.
Here are the details of the available versions:
* `jamba-large` currently points to `jamba-large-1.7-2025-07`
* `jamba-mini` currently points to `jamba-mini-2-2026-01`
* `jamba-large-1.7` points to `jamba-large-1.7-2025-07`
* `jamba-mini-2` points to `jamba-mini-2-2026-01`
## Model Deprecation
| Model | Snapshot | API Endpoint | Deprecation Date |
| --------------- | -------- | ------------------------- | ---------------- |
| Jamba Mini 1.7 | 2025-07 | `jamba-mini-1.7-2025-08` | 2026-02-01 |
| Jamba Large 1.6 | 2025-03 | `jamba-large-1.6-2025-03` | 2025-08-03 |
| Jamba Mini 1.6 | 2025-03 | `jamba-mini-1.6-2025-03` | 2025-08-03 |
| Jamba Large 1.5 | 2024-08 | `jamba-large-1.5-2024-08` | 2025-05-06 |
| Jamba Mini 1.5 | 2024-08 | `jamba-mini-1.5-2024-08` | 2025-05-06 |
## Model Compliance and Certifications
* [**SOC 2 compliance**](https://www.ai21.com/blog/soc-2-report)
* **ISO 27001, ISO 27017, and ISO 27018 certifications**
* [**Trust Center**](https://trust.ai21.com/?_gl=1*l2aasw*_gcl_aw*R0NMLjE3NDMzMTc1MTkuQ2owS0NRand0SjZfQmhEV0FSSXNBR2FubUtleDU5blZFaW9lRWctVG5HQno3Q092d2RoSUluVWpUdURPRWR1My1BdlBhRURGcWY0LVFzY2FBaFduRUFMd193Y0I.*_gcl_au*MTgyMTk2NzA0MS4xNzQyMjg0MjU3)
## Ethical Considerations
AI21 Labs is on a mission to supercharge human productivity with machines working alongside humans as thought partners, thereby promoting human welfare and prosperity. To deliver its promise, this technology must be deployed and used in a responsible and sustainable way, taking into consideration potential risks, including malicious use by bad actors, accidental misuse, and broader societal harms. We take these risks extremely seriously and put measures in place to mitigate them.
AI21 provides open access to Jamba that can be used to power a large variety of useful applications. We believe it is important to ensure that this technology is used in a responsible way, while allowing developers the freedom they need to experiment rapidly and deploy solutions at scale. Overall, we view the safe implementation of this technology as a partnership and collaboration between AI21 and our customers and encourage engagement and dialogue to raise the bar on responsible usage.
In order to use Jamba, you are required to comply with our [Terms of Service](https://lp.ai21.com/hubfs/resources/AI21-Models-Terms-of-Service.pdf?_gl=1*jr1jnx*_gcl_au*MTg4MDYzNjU4MC4xNzM0NTA4MjE4) and with the following [Usage Guidelines](/docs/responsible-use-1#usage-guidelines).
Please check these usage guidelines periodically, as they may be updated from time to time. For any questions, clarifications or concerns, please contact [safety@ai21.com.](mailto:safety@ai21.com)
## Limitations
There are a number of limitations inherent to neural networks technology that apply to Jamba. These limitations require explanation and carry important caveats for the application and usage of Jamba.
* **Accuracy:** Jamba, like other large pretrained language models, lacks important context about the world because it is trained on textual data and is not grounded in other modalities of experience such as video, real-world physical interaction, and human feedback. Like all language models, Jamba is far more accurate when responding to inputs similar to its training datasets. Novel inputs tend to generate higher variance in its output.
* **Coherence and consistency:** Responses from Jamba are sometimes inconsistent, contradictory, or contain seemingly random sentences and paragraphs.
* **Western/English bias:** Jamba is trained primarily on English-language text from the internet, and is best suited to classifying, searching, summarizing, and generating English text. Furthermore, Jamba has a tendency to hold and amplify the biases contained in its training dataset. As a result, groups of people who were not involved in the creation of the training data can be underrepresented, and stereotypes and prejudices can be perpetuated. Racial, religious, gender, socioeconomic, and other categorizations of human groups can be considered among these factors.
* **Explainability:** It is difficult to explain or predict how Jamba will respond without additional training and fine tuning. This is a common issue with neural networks of this scope and scale.
* **Recency:** Jamba was trained on a dataset created in March 2024, and therefore has no knowledge of events that have occurred after that date. We update our models regularly to keep them as current as possible, but there are notable gaps and inaccuracies in responses as a result of this lack of recency.
**Ready to take your projects to the next level?**\
Access Jamba via the [AI21 Studio API](/reference/jamba-1-6-api-ref) or [deploy privately](/docs/self-deployment).\
See our [Quick Start Guide](/docs/sdk) for integration steps, SDKs, and sample code.\
Explore Jamba in action on the [AI21 Studio Playground](https://studio.ai21.com/v2?tab=jamba_playground).
# Local Inference
Source: https://docs.ai21.com/docs/local-inference
For fully local execution, [llama.cpp](https://github.com/ggml-org/llama.cpp) enables running compatible open models in the GGUF format, with optional GPU acceleration. \
\
AI21 publishes official Jamba model weights on the [Hugging Face Hub](https://huggingface.co/ai21labs), and community contributors may provide GGUF-format conversions (e.g., Jamba Mini 1.7) for use with `llama.cpp`.
**Note:**
AI21 does not distribute or support GGUF builds and cannot verify the accuracy of third-party conversions.
Be sure to review the model's license terms and consult the `llama.cpp` documentation before use.
# Overview
Source: https://docs.ai21.com/docs/maestro-overview
## Introducing AI21 Maestro
AI21 Maestro is an AI system for rapidly creating and deploying RAG agents that automate high-value, data-intensive business tasks. At the core of AI21 Maestro is a new type of agent intelligence, optimized to find the smartest way to search, reason, validate, and adapt in real time to accomplish tasks, while staying within your cost and latency requirements.
## Key Benefits
* **Reliable Results with Built-In Validation** \
AI21 Maestro delivers accurate, high-quality outputs by selecting optimal tools, scaling compute resources as needed, and rigorously validating each step, all within your latency and cost constraints.
* **Scalable and Fast to Deploy** AI21 Maestro reduces time-to-value by automatically creating tailored execution plans. Simply define your goals, connect tools, and set your budget, AI21 Maestro handles the rest.
* **Full Transparency and Traceability** Every result includes an execution trace and a structured validation report, showing exactly how the system performed against your stated requirements.
## Product Details
At its core, AI21 Maestro is a dynamic planning system that determines the optimal sequence of actions to solve a given task during inference time. The system excels at self-validation and correction, continuously evaluating outputs against your specified requirements.\
\
Essentially, each call to AI21 Maestro builds a tree of calls to LLMs and other tools. \
Based on the task requirements and available budget, AI21 Maestro strategically plans which techniques to employ.
## Model Selection
AI21 Maestro is model-agnostic, it can orchestrate tasks using AI21’s first-party models or third-party models hosted by other providers. \
\
You can specify which model to use for a given run, or let Maestro automatically select the most suitable one based on your requirements.\
\
This flexibility lets you balance performance, latency, and cost while maintaining a unified interface. Available model types:
* **First-Party Models (AI21-hosted)**\
Optimized for reasoning and retrieval tasks, managed directly by AI21.\
Examples: `jamba-large`, `jamba-mini`.
* **Third-Party Models (Managed by AI21)**\
Access popular external models (e.g., OpenAI, Anthropic, Google) directly through the AI21 API, no extra setup required.\
Examples: `gpt-4.1`, `claude-4-sonnet`, `gemini-2.5-flash`, `mistral-7b` , `mistral-7b`.
* **Third-Party Models (BYOK)**\
Use your own API keys to access external models securely through Maestro.\
Supported providers include OpenAI, Anthropic, and Google.\
Configure BYOK models on the Third-Party Models page, then reference their IDs in your requests.
## Budget Control
The `budget` parameter lets you control that balance between **speed, cost, and reliability**.
A higher budget allows Maestro to explore more reasoning paths, take multiple execution steps, and validate results more thoroughly which can improve accuracy, but also increases latency and cost. A lower budget returns results faster and with lower compute usage, making it ideal for simpler or time-sensitive tasks.
**Budget levels:**
* `low`: Fastest and most cost-efficient. One execution step is taken with minimal effort.
* `medium`: Balanced performance for typical use cases. Multiple execution steps with moderate effort.
* `high`: Maximum reliability for complex or high-stakes tasks. Multiple strategies and validation cycles are applied.
Defaults to \_`low`*if not specified.*
## Response Language
AI21 Maestro can return outputs in multiple languages.\
You can control the response language for each run using the `response_language` parameter.\
Supported languages include **Arabic**, **Dutch**, **English**, **French**, **German**, **Hebrew**, **Italian**, **Portuguese**, and **Spanish**.\
If not specified, Maestro defaults to **English**.
## **Saving and Reusing Agents**
You can use AI21 Maestro to save agents and quickly reuse them in future API calls without redefining their configuration. A saved agent stores its name, instructions, tools, and configuration, so you don’t need to redefine them each time you invoke it.\
This makes it easy to keep your workflows consistent and efficient.
## Supported Use Cases
**Deep research agents for high-stakes tasks:**
* Financial report generation
* RFP response generation
* High-CapEx equipment troubleshooting
* M\&A due diligence
* Organization compliance review
* Contract portfolio analysis
**Complex Document Analysis:**
* Financial document summarization
* Investment prospectus analysis
* Clinical trial results analysis
* Technical documentation comparison
* Loan application evaluation
* Insurance claim analysis
**High-accuracy information parsing and extraction:**
* Legacy systems data migration
* Customer interactions intelligence
* Medical history encoding
* Supply chain data standardization
* Clinical trial results analysis
* Patent claim element extraction
* Contract term extraction
# RAG Overview
Source: https://docs.ai21.com/docs/maestro-rag-overview
Built on top of AI21’s advanced RAG Engine and enhanced with a planning layer, AI21 Maestro enables users to ask follow-up questions and receive grounded, context-aware answers.\
You can extend AI21 Maestro capabilities using built-in tools that provide access to additional context and information from the web or your files.
* **File Search:** Retrieve information from your uploaded documents.
* **Web Search:** Incorporate data from the web.
## Key Benefits
**Chat with Your Data**\
Go beyond single-question answering by allowing follow-up questions, clarifications, and step-by-step problem-solving through a natural conversation.
**Enterprise-Grade Accuracy**\
Built on retrieval-augmented generation (RAG), responses are grounded in your actual documents, not just model hallucination. It’s useful for internal tools and customer-facing applications alike.
**Fully Managed & Easy to Deploy**\
Simply upload your documents (PDF, DOCX, TXT, HTML, or Markdown). The RAG Engine automatically indexes them, making setup fast and seamless.
## **Deployment & Document Ingestion**
* Upload supported document types: .pdf, .docx, .txt, .md, .html.
* Indexing is automatic.
* The AI21 Maestro RAG system includes our in-house document parser, which provides high-quality parsing.
* Data connectors to cloud sources.
# RAG Quick Start
Source: https://docs.ai21.com/docs/maestro-rag-quickstart
RAG tools enable AI21 Maestro to access additional context, either from the AI21 Maestro File Library you manage or directly from the web. This helps generate grounded, relevant responses based on your data or real-time information.
## Available Tools
* **File Search** – Retrieve information from your uploaded documents.
* **Web Search** – Incorporate data from the web.
## Using File Search
Before using file\_search:
Upload your documents to your [File Library](https://docs.ai21.com/reference/manage-library-ref).
## **Using Web Search**
You can also enable `web_search` to let Maestro retrieve real-time information from the web.\
Optionally, you can restrict searches to specific domains using the `urls` parameter.
### Example:
The example below shows a completed `run` object that includes both `file_search` and `web_search` results. These `data_sources` fields only appear if you explicitly request them using the `include` parameter.
For details on how to use this parameter, see the `include` [parameter documentation](https://docs.ai21.com/reference/maestro-create-run#param-include).
```markdown theme={"system"}
{
"id": "dba286ef-5067-4c6e-b215-5483500da8df",
"status": "completed",
"result": "- example 1\n- example 2\n- example 3",
"data_sources": {
"web_search": [
{
"id": "0686b7a3-6cb6-7128-8000-146778a91763",
"text": "web search result text",
"url": "https://example.com",
"score": 0.01461567
}
],
"file_search": [
{
"text": "file search result text",
"file_id": "aa50de3e-1624-4e9f-a6f4-488a1aa16759",
"file_name": "example.pdf",
"score": 0.7331293
}
]
}
}
```
# RAG Use Cases
Source: https://docs.ai21.com/docs/maestro-rag-use-cases
Practical use cases where RAG enhances accuracy and relevance.
### Initialize the Client
```python theme={"system"}
from ai21 import AI21Client
client = AI21Client(api_key="YOUR_API_KEY")
```
### **Technical Troubleshooting Assistant**
**Challenge:**\
Support engineers often receive error codes from customers and must quickly identify the root cause. They typically need to search across scattered manuals, outdated notes, or legacy documentation, which is a slow process that delays issue resolution.
#### **Baseline Chat Model (Without RAG)**
When the model is prompted without access to contextual documents, it produces a general or incomplete answer, often missing key technical details.
**Input Prompt**
```python theme={"system"}
response = client.beta.maestro.runs.create_and_poll(
model="jamba-large",
input=[
{
"role": "user",
"content": """After power-up, the display shows Ch and won’t allow operation. Explain the compressor preheating logic (2.5-hour heat, first-install requires 6 hours of power applied), the condition for Ch to clear, and what is or isn’t safe to bypass. Provide a customer-friendly ETA message."""
}
]
)
print(response.result)
```
**Chat Model Output**
```python theme={"system"}
Display 'Ch' indicates the compressor is in preheating mode to protect it from damage. Preheating lasts 2.5 hours under normal conditions and 6 hours for first-time installation (requires continuous power). The 'Ch' clears automatically once preheating completes; do not bypass this safety process.
```
⚠️ **Issues Identified:**
* Lacks source attribution or evidence
* Provides generic advice not aligned with the specific product version
* Risks outdated or inaccurate troubleshooting guidance
### AI21 Maestro with RAG + Requirements
Adding **requirements** further improves reliability and structure by guiding the model to format and qualify its answers based on internal policy.\
\
**Before running the example:**\
Download the reference manual used in this example:\
📄[ air\_conditioner\_troubleshooting.pdf](https://storage.cloud.google.com/ai21-studio-staging-website-assets/SYSTEM%20AIR%20CONDITIONER.pdf)\
\
Upload it to your **File Library** in Maestro. The document will be automatically indexed for File Search, allowing Maestro to retrieve the correct sections during troubleshooting.\
\
**Step 1: Upload a file (Python SDK)**
```python theme={"system"}
file_from_disk = client.library.files.create(
file_path="/path/to/your/local/system_air_conditioner.pdf", # Replace with your file path
labels=["technical", "manual"] # Example labels; can be any descriptive tags
)
```
**Input Prompt**
```python theme={"system"}
response = client.beta.maestro.runs.create_and_poll(
models=["jamba-large"],
input=[
{
"role": "user",
"content": "After power-up, the display shows Ch and won’t allow operation. Explain the compressor preheating logic (2.5-hour heat, first-install requires 6 hours of power applied), the condition for Ch to clear, and what is or isn’t safe to bypass. Provide a customer-friendly ETA message"
}
],
tools=[
{
"type": "file_search",
"labels": ["technical", "manual"]
}
],
requirements=[
{
"name": "GroundedAnswer",
"description": "Use data from the attached manual; quote code labels and thresholds.",
"is_mandatory": True
},
{
"name": "OrderedDiagnosis",
"description": "If multiple causes are possible, order them by likelihood and cross-check against sensor readings or phase conditions.",
"is_mandatory": False
},
{
"name": "SafetyFlagging",
"description": "Flag and clearly label actions restricted to qualified technicians.",
"is_mandatory": True
},
{
"name": "StructuredOutput",
"description": "End with a plain-language customer summary followed by a detailed technical checklist for field technicians.",
"is_mandatory": False
}
]
)
print(response.result)
```
### Using `data_sources`
When you include `data_sources`, you explicitly tell Maestro to include**the data sources in the output.**
```python theme={"system"}
response = client.beta.maestro.runs.create_and_poll(
models=["jamba-large"],
input=[
{
"role": "user",
"content": """After power-up, the display shows Ch and won’t allow operation. Explain the compressor preheating logic (2.5-hour heat, first-install requires 6 hours of power applied), the condition for Ch to clear, and what is or isn’t safe to bypass. Provide a customer-friendly ETA message"""
}
],
tools=[
{
"type": "file_search",
"labels": ["technical", "manual"]
}
],
requirements=[
{
"name": "GroundedAnswer",
"description": "Use data from the attached manual; quote code labels and thresholds.",
"is_mandatory": True
},
{
"name": "OrderedDiagnosis",
"description": "If multiple causes are possible, order them by likelihood and cross-check against sensor readings or phase conditions.",
"is_mandatory": False
},
{
"name": "SafetyFlagging",
"description": "Flag and clearly label actions restricted to qualified technicians.",
"is_mandatory": True
},
{
"name": "StructuredOutput",
"description": "End with a plain-language customer summary followed by a detailed technical checklist for field technicians.",
"is_mandatory": False
}
],
include=["data_sources"]
)
print(response.result)
```
**AI21 Maestro (RAG + Requirements) Output**
```python theme={"system"}
Answer In Context
Compressor Preheating Logic
Preheating Duration: The compressor preheating logic involves a 2.5-hour heating period. This is necessary when the system is powered up and the display shows "Ch". During this time, the crankcase heater (CCH) heats the compressor.
First Installation Requirement: For first-time installations, it is required to apply power for 6 hours before operation to ensure proper preheating.
Condition for Clearing "Ch": The "Ch" indicator will clear once the compressor has been sufficiently heated for 2.5 hours, allowing the system to operate normally.
Possible Causes and Checks
Cause 1: Insufficient Preheating: Most likely, the system hasn't completed the 2.5-hour preheating. Ensure the CCH is operational.
Cause 2: Sensor Malfunction: Check the temperature sensor readings to ensure they reflect actual conditions.
Cause 3: Phase Error: Verify that the phase conditions are correct, as incorrect phase can prevent operation.
Safety and Bypass Considerations
Bypassing Preheating: It is not safe to bypass the preheating process as it is crucial for the protection of the compressor. Bypassing could lead to compressor damage or reduced efficiency.
Restricted Actions: Any attempt to bypass or modify the preheating logic should be performed only by qualified technicians.
Customer-Friendly ETA Message
"Your air conditioning system is currently in a preheating phase to ensure optimal performance and longevity. This process typically takes about 2.5 hours. For new installations, please ensure the system has been powered for at least 6 hours before operation. Thank you for your patience."
Technical Checklist for Field Technicians
Verify Power Application: Ensure that power has been applied for the required duration (6 hours for first-time installations).
Check Crankcase Heater (CCH): Confirm that the CCH is functioning correctly and heating the compressor.
Monitor "Ch" Indicator: Ensure the "Ch" indicator clears after the preheating period.
Sensor Readings: Cross-check temperature sensor readings to confirm proper heating.
Phase Conditions: Verify that phase conditions are correct and there are no phase errors.
Qualified Technician Actions: Any adjustments or bypassing of the preheating logic should be performed by a qualified technician only
```
***
**✅ Final Outcome**
* Combines document grounding with operational requirements
* Produces structured, role-specific responses
* Balances customer communication and technician detail
***
### Requirements for Reproducing the Example
* Download and upload the[ 📄 air\_conditioner\_troubleshooting.pdf](https://storage.cloud.google.com/ai21-studio-staging-website-assets/SYSTEM%20AIR%20CONDITIONER.pdf) file to your **File Library**.
* Enable **File Search** in your Maestro configuration.
* Use the **Python SDK** for consistency with other examples.
* Ensure documents are up-to-date to maintain accuracy.
# Account
Source: https://docs.ai21.com/docs/managing-your-account
# Overview
An account in AI21 grants you access to the playground and API usage. New accounts get a \$10 credit to use over the next three months. After that, you must provide billing information, and will be billed for any additional usage at [these rates](/docs/usage-cost). If you expect high usage, you can [contact us for bulk pricing](https://www.ai21.com/contact-sales/).
**Quick facts:**
* An AI21 *account* represents an **organization**.
* Each organization can have one or more **members** (sometimes called user accounts) identified by email. Currently AI21 allows a given email address to belong to one organization only.
* Members have a profile and a **role**, which determines what they can see or manage in the account (not what they can use on AI21).
* To use our API you must have an API key, which is available once you sign in.
## Sign up for AI21
There are two ways to get access to AI21 resources:
* **Join an existing organization account:** An AI21 organization administrator can invite you to join their organization. You'll get an email; follow the link in the email to join that organization. This means all permissions and usage will be managed by, and billed to, that organization. Your account user name is the email address where you received the invitation.
* **Create an account from scratch:** Visit the [account creation page](https://studio.ai21.com/sign-up) and either provide an email address or sign in using one of the authentication providers listed. Creating a new account this way creates a new organization account in AI21, for which you are the administrator. You can then invite other members, manage the organization's account, change or monitor your billing plan, or perform other administrative functions. You will need to provide billing information when your introductory [usage credits](/docs/usage-cost) expire.
One AI21 user account per email address
Currently, AI21 supports only one user account per email address. This means that the same email address cannot belong to multiple organizations in AI21.
## Signing in
Sign in to AI21 using any authentication method supported by AI21 (we support several methods, including SSO, Google, GitHub, and name/password). You must sign in to use the AI21 playground.
If your organization has enabled SSO, they might require you to sign in using SSO. If so, you'll see a warning message if you use any other sign-in method.
## Access your account management center
The account management center shows information about you and your organization. This includes member lists, usage and billing, and payment plans.
There are two ways to reach your account management center:
* **From the web app playground:** Click your profile picture at the top of the playground, then click **My account**.
* **From the documentation:** Click **Account** at the top of the page
# Billing and usage history
Billing and usage for an organization is the sum of billing and usage of all its members. The default usage plan is pay as you go, but if you have greater usage or support needs you can contact our sales department to design a custom billing plan.
Account administrators can see and change their billing plans:
* From the [account management center](#access-your-account-management-center), click **Account > Billing & Plans**
Account administrators can see the past few bills for their account:
* From the [account management center](#access-your-account-management-center), click **Account > Billing & Plans** to see your bill history.
Account administrators can see the usage numbers and costs for the current billing period for their organization:
* From the [account management center](#access-your-account-management-center), click **Account > Model usage** to see your bill history.
# Administer your organization's account
Users are identified by email address. AI21 limits each user to membership in one organization. If you add the same user to multiple organizations, the user will be able to access only the most recent organization they were added to.
You can manage your organization's account from the [account management center](#access-your-account-management-center).
**Users can have one of the following roles:**
* **Member:** Can use the playground and have full API usage.
* **Admin:** Same as **member**, but can also invite or remove members and manage member roles, see usage and billing for the organization, and change billing plans.
**Administrators can change the role of any other organization member:**
1. From the [account management center](#access-your-account-management-center), click **Organization** to see your organization's member list.
2. Click the *more options* icon ( … ) next to the member to manage, then choose the new role to assign.
From the [account management center](#access-your-account-management-center), click **Organization** to see your organization's member list.
**To invite a new member**
1. From the [account management center](#access-your-account-management-center), Select **Organization** to see your organization's member list.
2. Click **Invite member** to invite new members.
3. When the user has joined, [specify the user's role](#manage-users-role).
**To remove a member**
1. From the [account management center](#access-your-account-management-center), Select **Organization** to see your organization's member list.
2. Click the *more* menu more\_vert next to the member to remove.
3. Click **Delete** to remove the member's access. Note that the user might have the ability to navigate some of the playground or management pages for up to an hour, but cannot make changes or access the API.
AI21 supports SSO for organizations. AI21 supports all major identity providers, including Okta, Google Workspaces, OneLogin, and many more, as well as multi-factor authentication for extra security.
If you are interested in enabling SSO for your company's AI21 account, please [contact us](https://forms.gle/Si4B1Vd7MAj8n7ZY7).
If you have SSO enabled for your organization, you can specify whether members *must* use SSO when signing in, or if members can use any supported method to sign in. When you request SSO for your organization, let us know whether SSO should be required.
# MCP Server Setup
Source: https://docs.ai21.com/docs/mcp-server-setup-guide
## Building an Employee Server with FastMCP
This guide walks you through creating a Python-based MCP server that represents an employee management system. Using AI21 Maestro, you will be able to ask questions about your departments and employees. We’ll expose it remotely with ngrok and call it using AI21 Maestro.
## Prerequisites
* Python 3.10+
* ngrok account (free tier works)
* AI21 Maestro platform access
## Step 1: Create the MCP Server
### 1.1 Install uv and initialize the project
`uv` is a modern, extremely fast Python package and project manager.\
First, let’s install `uv` and set up our Python project and environment:
```bash theme={"system"}
curl -LsSf https://astral.sh/uv/install.sh | sh
```
Now let’s initialize the project:
```bash theme={"system"}
# Create a new directory for our project
uv init employee-server
cd employee-server
# Create virtual environment and activate it
uv venv
source .venv/bin/activate
```
Install dependencies:
```bash theme={"system"}
uv add fastmcp
```
### 1.2 Create the Server File
In the `employee-server` directory, create a file named `employee_server.py`:
```python theme={"system"}
from fastmcp import FastMCP
from typing import Dict, Optional
# Initialize FastMCP server
mcp = FastMCP("Employee Knowledge Base")
# Sample employee data (in production, this would come from a database)
EMPLOYEE_DATA = {
"EMP001": {
"name": "Alice Johnson",
"department": "Engineering",
"role": "Senior Developer",
"salary": 120000,
"email": "alice.johnson@company.com",
"manager": "EMP005"
},
"EMP002": {
"name": "Bob Smith",
"department": "Sales",
"role": "Sales Manager",
"salary": 95000,
"email": "bob.smith@company.com",
"manager": "EMP006"
},
"EMP003": {
"name": "Carol White",
"department": "HR",
"role": "HR Specialist",
"salary": 75000,
"email": "carol.white@company.com",
"manager": "EMP007"
},
"EMP004": {
"name": "David Brown",
"department": "Engineering",
"role": "Junior Developer",
"salary": 80000,
"email": "david.brown@company.com",
"manager": "EMP005"
},
"EMP005": {
"name": "Eva Martinez",
"department": "Engineering",
"role": "Engineering Manager",
"salary": 150000,
"email": "eva.martinez@company.com",
"manager": "EMP008"
}
}
@mcp.tool()
def get_employee_by_id(employee_id: str) -> Dict:
"""
Retrieve employee information by their ID.
Args:
employee_id: The unique employee identifier (e.g., EMP001)
Returns:
Employee information including name, department, role, and salary
"""
if employee_id in EMPLOYEE_DATA:
return {
"success": True,
"data": EMPLOYEE_DATA[employee_id]
}
else:
return {
"success": False,
"error": f"Employee with ID {employee_id} not found"
}
@mcp.tool()
def search_employee_by_name(name_query: str, max_results: Optional[int] = 10) -> Dict:
"""
Search for employees by name using fuzzy matching.
Args:
name_query: The name or partial name to search for
max_results: Maximum number of results to return (default: 10)
Returns:
List of employees matching the name query, sorted by relevance
"""
if not name_query.strip():
return {
"success": False,
"error": "Name query cannot be empty"
}
query = name_query.lower().strip()
matches = []
for emp_id, emp_data in EMPLOYEE_DATA.items():
employee_name = emp_data["name"].lower()
score = 0
# Exact match gets highest score
if query == employee_name:
score = 100
# Check if query is contained in name
elif query in employee_name:
score = 80
# Check if all words in query are in name
elif all(word in employee_name for word in query.split()):
score = 60
# Check if any word in query matches any word in name
elif any(word in employee_name for word in query.split()):
score = 40
# Check if name starts with query
elif employee_name.startswith(query):
score = 70
# Check for partial word matches
else:
query_words = query.split()
name_words = employee_name.split()
partial_matches = 0
for q_word in query_words:
for n_word in name_words:
if q_word in n_word or n_word in q_word:
partial_matches += 1
break
if partial_matches > 0:
score = 20 + (partial_matches * 10)
if score > 0:
matches.append({
"id": emp_id,
"score": score
})
# Sort by score (highest first)
matches.sort(key=lambda x: x["score"], reverse=True)
# Limit results
if max_results:
matches = matches[:max_results]
return {
"success": True,
"query": name_query,
"count": len(matches),
"employee_ids": [match["id"] for match in matches]
}
@mcp.tool()
def search_employees_by_department(department: str) -> Dict:
"""
Search for all employees in a specific department.
Args:
department: The department name to search for
Returns:
List of employees in the specified department
"""
employees = []
for emp_id, emp_data in EMPLOYEE_DATA.items():
if emp_data["department"].lower() == department.lower():
employees.append({
"id": emp_id,
**emp_data
})
return {
"success": True,
"count": len(employees),
"employees": employees
}
@mcp.tool()
def get_salary_range(min_salary: Optional[int] = None, max_salary: Optional[int] = None) -> Dict:
"""
Find employees within a specific salary range.
Args:
min_salary: Minimum salary threshold (optional)
max_salary: Maximum salary threshold (optional)
Returns:
List of employees within the specified salary range
"""
employees = []
for emp_id, emp_data in EMPLOYEE_DATA.items():
salary = emp_data["salary"]
if (min_salary is None or salary >= min_salary) and \
(max_salary is None or salary <= max_salary):
employees.append({
"id": emp_id,
"name": emp_data["name"],
"department": emp_data["department"],
"salary": salary
})
# Sort by salary
employees.sort(key=lambda x: x["salary"], reverse=True)
return {
"success": True,
"count": len(employees),
"employees": employees
}
@mcp.tool()
def get_employee_hierarchy(employee_id: str) -> Dict:
"""
Get the reporting hierarchy for an employee.
Args:
employee_id: The employee ID to get hierarchy for
Returns:
The employee's manager and any direct reports
"""
if employee_id not in EMPLOYEE_DATA:
return {
"success": False,
"error": f"Employee with ID {employee_id} not found"
}
employee = EMPLOYEE_DATA[employee_id]
# Find manager
manager = None
if employee.get("manager") and employee["manager"] in EMPLOYEE_DATA:
manager = {
"id": employee["manager"],
"name": EMPLOYEE_DATA[employee["manager"]]["name"],
"role": EMPLOYEE_DATA[employee["manager"]]["role"]
}
# Find direct reports
direct_reports = []
for emp_id, emp_data in EMPLOYEE_DATA.items():
if emp_data.get("manager") == employee_id:
direct_reports.append({
"id": emp_id,
"name": emp_data["name"],
"role": emp_data["role"]
})
return {
"success": True,
"employee": {
"id": employee_id,
"name": employee["name"],
"role": employee["role"]
},
"manager": manager,
"direct_reports": direct_reports
}
@mcp.tool()
def get_department_statistics() -> Dict:
"""
Get statistics for all departments including employee count and average salary.
Returns:
Statistics for each department
"""
dept_stats = {}
for emp_data in EMPLOYEE_DATA.values():
dept = emp_data["department"]
if dept not in dept_stats:
dept_stats[dept] = {
"count": 0,
"total_salary": 0,
"employees": []
}
dept_stats[dept]["count"] += 1
dept_stats[dept]["total_salary"] += emp_data["salary"]
dept_stats[dept]["employees"].append(emp_data["name"])
# Calculate averages
result = {}
for dept, stats in dept_stats.items():
result[dept] = {
"employee_count": stats["count"],
"average_salary": stats["total_salary"] / stats["count"],
"total_salary": stats["total_salary"],
"employees": stats["employees"]
}
return {
"success": True,
"departments": result
}
# Run the server
if __name__ == "__main__":
mcp.run(transport="streamable-http", path="/mcp", port=8000)
```
### 1.3 Run the Local Server in development mode
```bash theme={"system"}
python employee_server.py
```
Your server should now be running at \[http\://localhost:8000.]\([http://localhost:8000.\\)\\\\](http://localhost:8000.\\\)\\\\) Test it by visiting [http://localhost:8000/mcp](http://localhost:8000/mcp) and getting the following response:
```bash theme={"system"}
{"jsonrpc":"2.0","id":"server-error","error":{"code":-32600,"message":"Not Acceptable: Client must accept text/event-stream"}}
```
###
## Step 2: Test the MCP Server
Now we will test that the server is running and that it adheres to the Model Context Protocol correctly by listing its tools.
We will use the [MCP Inspector](https://github.com/modelcontextprotocol/inspector) tool which creates a locally running MCP client along with a user facing web interface.
### 2.1 Run the MCP Inspector
In a different terminal run the following command:
```bash theme={"system"}
npx @modelcontextprotocol/inspector@latest
```
You should see the following output:
```bash theme={"system"}
⚙️ Proxy server listening on localhost:6277
🔑 Session token:
Use this token to authenticate requests or set DANGEROUSLY_OMIT_AUTH=true to disable auth
🚀 MCP Inspector is up and running at:
```
Copy the session token from this output.
### 2.2 List your server tools
Open the browser in `http://localhost:6274/`and you should see the Inspector UI.
1. Choose Streamable HTTP as the Transport Type
2. Set the URL to `http://localhost:8000/mcp`
3. Go to the **Configuration** pane and paste the `` in the Proxy Session Token textbox
4. Click on the **Connect** button
5. Select Tools and then click on “List Tools”
6. You should see a list of all the tools your server exposes
7. **Optional**: you can interact with each tool by clicking its name in the list and providing the required parameters
## Step 3: Expose Server with ngrok
### 2.1 Install ngrok
Download ngrok from [ngrok.com](https://ngrok.com/download) and create a free account.
### 2.2 Start ngrok Tunnel
```bash theme={"system"}
ngrok http 8000
```
You'll get a publicly accessible URL like: `https://abc123.ngrok-free.app`
**Important**: Save this URL, you'll need it for AI21 Maestro integration.
Test the URL by making a request to `https://abc123.ngrok-free.app/mcp` . and getting You should get the following response:
```bash theme={"system"}
{"jsonrpc":"2.0","id":"server-error","error":{"code":-32600,"message":"Not Acceptable: Client must accept text/event-stream"}}
```
## Step 4: Integrate with AI21 Maestro
### 4.1 Run with AI21 Maestro
In your AI21 Maestro configuration, add the remote MCP server:
```python theme={"system"}
import asyncio
from ai21 import AsyncAI21Client
client = AsyncAI21Client(api_key="")
async def main():
run = await client.beta.maestro.runs.create_and_poll(
input="Who reports to Eva Martinez?",
tools=[
{
"type": "mcp",
"server_url": "https://.ngrok-free.app/mcp",
"server_label": "Employees",
},
],
budget="medium",
)
print("id:", run.id)
print("Status:", run.status)
print("Result:", run.result)
# Works in Jupyter
await main()
# Comment the above and uncomment the below to run in Python scripts
# import asyncio
# asyncio.run(main())
```
You should get the following response:
```
id: 068c664b-f9ea-7078-8000-be3f80ffa83c
Result: Eva Martinez, who is an Engineering Manager, has the following direct reports:
1. Alice Johnson - Senior Developer
2. David Brown - Junior Developer
```
### 4.2 Test in AI21 Maestro
Example queries to test in AI21 Maestro:
1. "What is the salary of employee EMP001?"
2. "Show me all employees in the Engineering department"
3. "Who reports to EMP005?"
4. "What are the department statistics?"
# MCP
Source: https://docs.ai21.com/docs/mcp-studio-guide
How to Use MCP in AI21 Studio
MCP (Model Context Protocol) is an open-source standard for connecting AI applications to external systems. Follow these steps to connect your agent in **AI21 Studio** to an MCP server.
### Step 1: **Go to the Playground Section**
Navigate to Maestro [**Playground**](https://studio.ai21.com/v2?tab=maestro) from the left-hand menu in AI21 Studio.
### Step 2: Add a Tool
In the **Configuration** panel, go to the **Tools** section and click **Add**.
### **Step 3: Select MCP**
In the tools dropdown, select **MCP**.
### **Step 4: Fill in the Fields**
**4.1 Add a URL**\
Enter the MCP server address in the **URL** field.
**Handling “Invalid URL” Errors**\
\
If you see the **“Invalid URL”** error when adding your MCP server in AI21 Studio, it usually means:
* The server is using `http://` instead of `https://`
* The server is running on a raw IP address and isn’t exposed as a public HTTPS endpoint
**How to fix it:** Set up an **ngrok tunnel** (or a similar tool) to expose your local server over HTTPS, which passes validation.
\
For full setup instructions, see the [MCP Server Setup guide](https://docs.ai21.com/docs/mcp-server-setup-guide).
\
**4.2 Add a Label**\
Enter a descriptive **Label** to identify the MCP server.
**4.3 Add Authentication** \
From the **Authentication** dropdown, select the authentication method (for example, API key or access token) and provide the credentials.
Authorization headers are not saved with the agent. If you save the agent, the headers will not be stored. When you reload the agent, you’ll need to enter the authorization headers again.
### **Step 5: Review Allowed Tools**
Click **Review allowed tools**.
You will see a list of tools available from the MCP server.
### **Step 6: Save the Tools**
After verifying the tools are correct, click **Save**.
### **Step 7 (Optional): Edit MCP Tools**
Edit an existing MCP tool by following these steps:
1. Click the **⋮** (colon icon).
2. Select **Edit**, make your changes, and click **Save.**
### Next Steps
Use MCP in [AI21 Stuido](https://studio.ai21.com/v2?tab=maestro).
# Model Availability by Platform
Source: https://docs.ai21.com/docs/model-availability-across-platforms
## Where to Find AI21's Jamba Models
AI21's Jamba models are available across multiple leading cloud platforms and model hubs. Choose the platform that best fits your infrastructure and deployment needs.
## Quick Platform Overview
***
## HuggingFace
**Platform:** HuggingFace\
**Best for:** Research, open-source projects, local development, fine-tuning
**Platform:** HuggingFace\
**Best for:** Research, open-source projects, local development, fine-tuning
**Platform:** HuggingFace\
**Best for:** Research, open-source projects, local development, fine-tuning
***
## Kaggle
**Platform:** Kaggle\
**Best for:** Research, open-source projects, local development, fine-tuning
***
## Google Cloud Platform (GCP)
### Self-Deployment
**Platform:** GCP Model Garden (Self-Deploy)\
**Best for:** ML Engineers, Custom infrastructure, on-premises deployment
### Coming Soon
**Platform:** GCP Model Garden (Self-Deploy)\
**Status:** Coming Soon 🔄
***
## Microsoft
### Self-Deployment
**Platform:** Microsoft Foundry (Self-Deploy)\
**Best for:** ML Engineers, Custom infrastructure, on-premises deployment
***
## Amazon Web Services (AWS)
### Managed Services
**Platform:** AWS Bedrock (Managed)\
**Best for:** Enterprise AWS users, serverless applications, pay-per-use pricing
**Platform:** AWS Bedrock (Managed)\
**Best for:** Enterprise AWS users, serverless applications, pay-per-use pricing
### Self-Deployment
**Platform:** AWS SageMaker\
**Best for:** ML Engineers, Custom infrastructure, on-premises deployment
**Platform:** AWS SageMaker\
**Best for:** ML Engineers, Custom infrastructure, on-premises deployment
### Coming Soon
**Platform:** AWS Bedrock\
**Status:** Coming Soon 🔄
**Platform:** AWS Bedrock\
**Status:** Coming Soon 🔄
***
## Interested in Self-Deploy?
For self-deployment on your own infrastructure, check out our [vLLM deployment guide](/docs/vllm) for step-by-step instructions and examples.
# Overview
Source: https://docs.ai21.com/docs/overview
## Welcome to the AI21 Developer Platform!
AI21 provides AI systems and foundation models designed to address high-value, data-intensive workflows and real-world challenges. Our solutions are reliable, efficient, and transparent, especially effective for long-context tasks critical to enterprises, such as:
* Grounded question answering across lengthy documents
* Chat completion
* Financial data analysis
* Retrieval-Augmented Generation (RAG) workflows
## AI21 Maestro
AI21 Maestro is an advanced AI system for rapidly creating and deploying **knowledge agents** that automate high-value, data-intensive business tasks. It includes RAG capabilities, allowing models to retrieve information from a knowledge base of previously uploaded files using semantic search and web search.
At its core, AI21 Maestro features a new type of agent intelligence optimized to search, reason, validate, and adapt in real time, while staying within your cost and latency requirements. The system excels at self-validation and correction, continuously evaluating outputs against your specified goals.
Learn more about [AI21 Maestro](/docs/maestro-overview).
## Jamba Family of Open Models
Jamba models deliver up to **2.5× faster** inference than leading models of similar size. They are available for **private deployment**, including VPC and on-premises, and are optimized for tasks like:
* Long-context RAG
* Grounded QA
* Data classification
Jamba offers tailored performance for productivity tools, customer service, and internal knowledge agents. Learn more about the [Jamba Models](/docs/jamba-foundation-models).
***
## Flexible Deployments for the Enterprise
You can choose the option that meets your needs [here](https://www.ai21.com/deployment/).
## Accessing our Models and Solutions
You can access our tools and models through several mechanisms:
* **Python SDK:** We provide a [Python SDK](/docs/sdk) to simplify access to all our models and tools from your Python code. The SDK provides code completion tips, documentation, support for synchronous and asynchronous calls, and much more.
***
* **REST API:** Under the hood, our SDK, playground, and cloud platform implementations access our models through our [public REST API](/reference/jamba-1-6-api-ref).
***
* **Cloud platforms:** We provide model access from all the leading cloud providers. See the full list [here](/docs/model-availability-across-platforms).
***
* **Other third-party services:** Our models are also available on other third-party systems such as [LangChain](https://python.langchain.com/docs/integrations/chat/ai21/) and [LlamaIndex](https://docs.llamaindex.ai/en/stable/examples/llm/ai21/). Check your toolchain to see if AI21 is available.
## Next steps
* [Learn how to use our SDK](/docs/sdk).
***
# Prompt Engineering for Jamba Models
Source: https://docs.ai21.com/docs/prompt-engineering
### Overview
Prompt engineering is the practice of creating the proper prompt to generate the output that you want. With the proper prompting, an LLM can do an amazing number of tasks, including generating an email or product description, summarizing provided text, classifying text into standard or customized categories, responding to a customer query with an appropriate answer, and much more.
Every model behaves differently given the same prompt, and so you'll probably spend a fair amount of time adjusting your prompt for your specific use case or adapting it for different models. The best practices given here are not absolute rules, but guidelines.
***
### Components of a prompt
A prompt, in our context, consists of the following information. Other than the instruction, all other elements depend on the task and improving or customizing the response.
* Instruction: What the model should do, often in great detail.
* The format or syntax of the response \
(Markdown/text, HTML, JSON, XML).
* The persona of the LLM \
("a friendly travel agent").
* Any stylistic or other guidelines for the output \
("simple, plain language, no more than two sentences long")
* Examples to follow, if appropriate.
* Any text to be analyzed or acted upon \
(such as a user's question or financial or medical information)
***
### Developing the prompt
Your first prompt is rarely good enough, particularly when designing a prompt to use for a commercial system. You'll spend a lot of time refining your prompt and assessing the results.
**Here is a typical workflow for designing and refining a prompt:**
1. Start with just the instruction, without examples. You can ask the LLM for help creating a starter prompt and modify it from there.
2. Test your prompt against one example, and evaluate the quality.
3. Modify the prompt, test, and repeat.
* While you're adjusting your prompt, set the temperature to 0 to get consistent answers, then gradually increase the temperature in small increments up to where you are getting the amount of variation that you want.
* When the model produces bad output, bring the temperature back to zero and adjust the model again to try to understand what's producing the bad output.
4. Collect a test set of 10 common inputs and ideal outputs and test your prompt against those inputs.
* Be sure that your test prompts include both common inputs and inputs that test the extremes of valid requests.
* Include test examples of bad input to see how the model reacts to it.
5. Grade the test results against the expected result (you can grade on a score from 1-10 or 1-100). Evaluate the results, adjust your prompt, and repeat.
***
## Prompt design tips
**Be concise**
Say everything you want to say with as few words as possible. Don't state the obvious.
You are a customer support representative for ACME corp, your name is Wile E. Coyote. you need to answer user questions regarding support issues. Be polite, engaging and to the point.
Do not curse.
Do not mention competitors.
You are Wile E., ACME corp support chat representative.
Be polite, engaging and to the point.
**Ensure that the prompt is clear**
When crafting prompts for Jamba models, follow this fundamental principle: Write your prompts as you would want them to be written for you. If you can't understand the prompt well, neither will an LLM.
**To elaborate:**
* **Clarity and Comprehensibility:** Ensure that your prompt is free from grammatical errors and that the instructions are expressed clearly and unambiguously.
* **Formatting and Examples:** Having a structured output allows further manipulation after generation.
List all the following animals, objects and places in the story.
Story:
Your lists:
List all the following animals, objects and places in the story.\
Your output should be in the following format:
```json theme={"system"}
{
"animals": ["all animals in the story"],
"objects": ["all objects in the story"],
"places": ["all places in the story"]
}
```
Story: story
– End of story –
Clean and well-structured prompts minimize errors and help you debug and optimize your instructions. If you encounter difficulties or errors that you can't seem to fix, try simplifying your prompt.
**Describe the DOs, not the DON'Ts**
Focus on telling the model what you want to do. Minimize "do nots." Of course, you can use negatives occasionally, but excessive use of "don't do" and "avoid x" is a sign that your **prompt may be going in the wrong direction** and you may want to rewrite your prompt.
Write a product description for a high-end cell phone (i.e. not a landline). The description
should not be for regular folks; it should only be for important executives.
Do not make it overly sales-ish; instead have it be grounded in the specs of the phone.
Write a product description for a high-end cell phone. The description should be tailored
for a high powered executive and focus on the specs of the phone. Focus on how the phone can enable more efficient work.
**Allow the model to say "I don't know"**
Specifically state to the model you permit it to not return an answer. This reduces hallucinations.
Why did revenue increase, according to the following quarterly report?
Quarterly Report
Why did revenue increase, according to the following quarterly report?
If the answer is not in the provided report, reply only with "I don't know"
Quarterly Report
**System prompt - Use it!**
Describe the role that the LLM should assume when answering the question. This is frequently referred to as a *system prompt*, and has been found to produce better completions for many families of LLMs. This affects not only the tone and language used, but also the amount of detail and level of expertise used.
System prompts should also be used to guide the perspective the model takes when answering the question; for example, by thinking about the problem as a research assistant, or a customer, or a novice.
When accessing the model in code, the system prompt is specified by an initial `role:system` message. In the playground, provide the information in the *System instructions* section. Alternatively, you can put the system prompt directly into the prompt itself, although this might be less effective.
I want you to assume the role of a meticulous research assistant.\
Your task is to Evaluate if the following text extract is relevant to the case at hand.
You are a meticulous, critical research assistant.
```python In the SDK theme={"system"}
def single_message_instruct():
prompt = """ Evaluate if the following text extract is relevant to the case at hand.
Extract:
...
Case:
...
"""
messages = [
ChatMessage(
role="system",
content="You are a meticulous research assistant"
),
ChatMessage(
role="user",
content=prompt
)
]
response = client.chat.completions.create(
model="jamba-instruct",
messages=messages,
temperature=0.7
)
print(response.choices[0].message.content)
```
**The IDH Template (Instruction—Data—Hint)**\
For complex prompts, include instructions, then data, then a hint. This can be used for simple prompts as well.
The hint should be a paraphrased version of the instructions.
Rewrite the following patient record, so that it is easily understandable for an average person with a high school degree.
```
{PATIENT RECORD}
```
Rewrite the following patient record, so that it is easily understandable for an average person with a high school degree.
```
{PATIENT RECORD}
```
high school level rewrite:
For prompts with a lot of data it is better to clearly state to the model where every section starts and ends.
Your task is to fix the product description to be compliant with the product guidance.
`{Product Description}`
`{Product Guidance}`
Your product description:
Your task is to fix the product description to be compliant with the product guidance.
Original product description:
`{Product Description}`
– End of Description –
Product Guidance:
`{Product Guidance}`
– End of Description –
Your product description:
For some prompts with complex instructions, it is useful to include instruction both in the beginning and in the end of the prompt.
Your task is to fix the product description to be compliant with the product guidance.
Original product description:
`{Product Description}`
– End of Description –
Product Guidance:
`{VERY COMPLEX Product Guidance}`
– End of Description –
Your product description:
Your task is to fix the product description to be compliant with the product guidance.
Original product description:
`{Product Description}`
– End of Description –
Product Guidance:
`{VERY COMPLEX Product Guidance}`
– End of Description –
The rewrite product description in accordance with the product guidance:
**Use structured output when needed**If your output is meant to be read by another system (e.g., for integration into a pipeline), request a JSON-formatted response.\
Use the response\_format=json API parameter and specify the expected structure in the prompt itself.\
\
**Maintaining output quality in production**
Log your prompts and responses in production. Periodically check the output quality for a random selection of your generated answers.Provide a feedback button that sends you the question and generated response to help you improve your prompts.\
\
**Use structured output when needed**I\
If your output is meant to be read by another system (e.g., for integration into a pipeline), request a JSON-formatted response.\
Use the response\_format=json API parameter and specify the expected structure in the prompt itself.
Extract the user’s name, location, and request from the input text.
Extract the user’s name, location, and request from the input text.
Return the output in the following JSON format:
`{ "name": "", "location": "", "request": "" }`
**Use the appropriate tool**\
For straightforward math calculations or other actions that can be done in simple code, use a more appropriate tool for the job (a calculator, a macro, a short code snippet). Those tools are designed specifically for the job, and provide much more controllable and consistent output than an LLM.\
\
**Ask the LLM if it understands the prompt**\
To speed up development, consider first asking the LLM if it understands the instructions and other key terms in the prompt. This can help ensure that the LLM understands the core idea of what you are trying to do.\
\
**Don’t do math**\
LLMs are famously bad at math. They’re getting better, but it’s still not a good idea, generally to ask an LLM to solve a word problem for you. LLMs can count and do sums, but word problems or logic are tricky and produce inconsistent results. For simple or straightforward math, use a calculator or other more appropriate tool.If you must evaluate a word problem or logic, you can increase the accuracy by asking the LLM to explain each step it takes in the process (called *chain-of-thought prompting*.\
\
**Classification/Ranking**\
When performing classification or other scoring, it is much better to use categories that are meaningful rather than arbitrary numbers.
Analyze whether the two sentences provided are consistent or inconsistent. You can provide a score of between 1-3. \
Sentence A:
“When talking on the phone, the defendant confessed to felony murder.”
\
Sentence B:
“The defendant admitted nothing when talking on the phone.”
Your score:
Analyze whether the two sentences provided are consistent or inconsistent. Classify into one of the following classes: \[Consistent, Partially Consistent, Inconsistent] Sentence A:
“When talking on the phone, the defendant confessed to felony murder.”
\
Sentence B:
“The defendant admitted nothing when talking on the phone.”
Your classification:
## **Consistency vs. creativity**
Output variability is influenced by two parameters:*temperature\_and\_top P*:
* *Temperature* Flattens or enhances the curve of probability values for all candidates in the pool. When the curve is completely flat (T=max) there is an equal likelihood for any candidate in the pool to be chosen, no matter how unlikely the candidate originally was. When the curve is completely enhanced (T=0), the most likely candidate is now the only choice, and all the other candidates have probability zero. Temperature does not adjust the size of the pool by eliminating candidates, unless you choose T=0, which essentially removes all candidates from the pool except the most likely. For most use cases, a temperature of around 0.7 is about right.
* *Top P* controls the size of the pool of candidates for the next token (remember P for Pool). The smaller the candidate pool, the less variability. Candidates are cut in order of least likely to most likely, so at its smallest size (0.01), only the most likely candidates–possibly only one–will be in the pool. At 1.0, the pool is full sized and contains all possible candidates (including some possible weirdos). In practice you won’t need to adjust the top P often.
If you need completely consistent answers, such as for classification or math (but [**don’t do math**](https://docs.ai21.com/docs/prompt-engineering#dont-do-math)), set the temperature to 0.When you want some variation, start with a low temperature (0.2-0.3) and increase by tenths until you get the variability that you want. Typically you won’t need a temperature higher than 0.7. Note that setting temperatures higher than 0.7 can cause the LLM to wander, and setting it higher than 1.0 can cause extremely long and sometimes nonsensical output (definitely specify a`max_tokens`limit for high temperatures).If the temperature is very high (greater than 0.7) reduce the top P a few tenths of a point to remove the very unlikely results unless you want some very high creativity, in which case you can keep the topP high (try a topP of 0.99 to omit extremes).\
\
**Limiting output length**\
If you have a length goal or limit for your output, specify it in the prompt as the desired or maximum number of lines, words, sentences, or paragraphs. Don’t expect the model to hit the mark exactly, as it can only see one word ahead, and it might need a bit more or less than you specify to provide a good answer.\
**Examples:**“Write two or three sentences about…,” or “Limit your answer to 10 words.”The API supports a`max_tokens`parameter, but you should use it only as a failsafe to prevent the edge case of the model going off in a completely unexpected direction and far exceeding your length limits (higher temperature can increase output length). This value is absolutely respected, so if the output hits this limit, the result might stop in the middle of a word. Set this value a fair bit higher than any prompt-suggested limit.\
\
**Requesting citations**\
LLMs have been known to invent citations, so asking Jamba Instruct for citations for its information is not a guarantee of accuracy. If you need absolutely reliable citations, use the[**RAG Engine**](https://docs.ai21.com/v4.1/docs/rag-engine-overview).\
\
**Use labels, not numbers, to rate output**\
It is common for people to want a numerical scoring of “good” or “bad” (“on a scale of 1 to 10…”). Assigning exact numbers to subjective categories is hard for people, and harder for LLMs, and also gives a false sense of accuracy. Simple labels like “Bad,” “Okay,” and “Best” are easier for the language model to provide than precise numerical ratings like 4.7 or 6.8.
Although you can use numbers to represent categories it is generally preferable to use category labels that are inherently meaningful, such as “None”, “Some”, and “Most”.
```csharp theme={"system"}
I want you to analyze whether the two sentences provided are consistent or inconsistent.
You can provide a score of None, Some, Most, where “Most” means the most inconsistent.
#####
Sentence A: Josh went to the store and then went to school
Sentence B: Josh never went to school
Output: Most
#####
Sentence A: Josh has a dog and a cat
Sentence B: Josh has a fish
Output: None
####
Sentence A: Josh prefers typing on his laptop
Sentence B: Josh prefers typing on his phone and his laptop
Output: Some
####
Sentence A: When talking on the phone, the defendant confessed to felony murder.
Sentence B: The defendant admitted nothing when talking on the phone.
```
### **Providing examples**
Providing examples of what you want to see (also called “Few Shot prompts”) can be few useful to Jamba. Examples can be very helpful when
1. Providing clear instruction is laborious/difficult, or
2. The model seems to not fully understand your instructions.
Before you provide examples, try out the results using just instruction to see if you get the results you need. If it turns out that the results look better with an example or two, or if describing something is harder than showing an example, then go ahead and use examples.Some recommendations about using examples in your prompt:
* Separate examples clearly with a blank line, or another obvious marker (some people use ###).
* Put the instructions all at the beginning or at the end.
* Be sure all examples use the same structure; don’t show different examples with different fields, or totally different structure.
* When providing just a few examples, the ordering of examples probably doesn’t matter. When using the RAG engine, or providing very, very large examples or input, you might find that the order of files used as input might affect the output (a “recency bias”).
In the following example, we use two types of delimiters: a delimiter ### between examples and a newline between ads and answers. We’ve put the examples first and the instructions at the end. Content between user\[” ”] marks are just placeholders in the example where you would put actual ads, answers, or criteria.
```csharp theme={"system"}
Examples:
EXAMPLE AD 1
EXAMPLE AD 2
EXAMPLE ANSWER 1
####
EXAMPLE AD 3
EXAMPLE AD 4
EXAMPLE ANSWER 2
###
I want you to decide which of the following two advertisements are best, given the following criteria.
Criteria:
- Shorter is generally better
- If there is some kind of rhyme within the ad, that can add a little value
- References to superheroes are a huge plus
...
AD 1
AD 2
```
**Cover all cases in your examples**\
When providing examples, do your best to cover all relevant example types. For example, if the answer to a question posed to the LLM can be “yes” or “no,” provide examples where the answer is yes and examples where the answer is no. Ideally, the distribution should match real-life use cases as [well.In](http://well.In) the following example, we provide prompt examples that show how to respond to both vague and specific feedback from the user.
```swift theme={"system"}
For the following negative review of a product, suggest a way the comment can be addressed by a
product change if the review is negative. If the review contains nothing specific, pen a friendly
message asking for clarification for what is wrong.
####
Input:
I really didn't like your blender, it was defective.
Output:
Hi! We really appreciate your feedback about the blender. Can you elaborate what specifically
was defective? We are always improving our blenders and really would like to hear what went wrong.
####
Input:
I recently bought your top of the line laptop with 64K memory, but it ran out of memory.
Output:
We should consider increasing the hard drive storage to be significantly larger than 64 GB.
####
Input: I recently bought a car from this company and was extremely disappointed to find that
it broke down multiple times within the first week. The car had a number of issues that should
have been fixed before being sold.
Output:
```
**Grading results**
During development, and also after release (if you’re using your prompt in a production system) you’ll need to evaluate your results. Which methods you use depend on your usage scale and where you are in the development process.For grading outputs with an absolute answer (such as classification or sentiment analysis), judge on the following criteria:
* Accuracy (how many times the answer was correct)
* For prompts that are sensitive to wrong answers, calculate false positives and false negatives, and the cost of each, when deciding the bar to reach before deciding that your prompt is good enough.
Create an ideal answer for each test input and grade the generated result against the golden answer on a scale of 1-10. This is especially useful for classification exercises, which you will mark as correct or incorrect.\
\
**Human evaluation**\
If you have the resources, human evaluation is often the best solution. Typically in the early stages of adjusting your prompts you’ll be evaluating the answers yourself. Create a list of criteria to consider when evaluating each answer (accuracy, clarity, usefulness). You can either rate each criteria individually or give a general overall score (1-5, good/OK/bad).\
\
**Use an LLM to evaluate your results**\
You can try using another LLM to rank your answers. If you do, use a different LLM than the one that generated the answer (LLMs, like people, tend to be biased toward their own answers). For example, if you have a large block of information and your prompt asks the model to answer a specific question based on that information, you might write a prompt like this to evaluate the result generated from your first prompt:
```
<>
###
Based on the context above, does the following answer provide a correct and full answer to the question based on this context.
Pay attention to the following criteria:
- The answer must be correct according to the information in the context
- The answer that uses information that is not available in the context is a bad answer.
- If the question can’t be answered by the information in the context, the answer should be “Answer not in document
- If the answer is in the context, and is provided, that should increase the score of the answer.
- The answer must fully answer the question and not be too short.
- All relevant information should be in the answer.
- A vague or general answer is not very good.
Reply in the following format:
- Score: Good, OK, or Bad
- Reasoning: Why this score was chosen.
###
<>
```
**Maintaining output quality in production**
Log your prompts and responses in production. Periodically check the output quality for a random selection of your generated answers.Provide a feedback button that sends you the question and generated response to help you improve your prompts.\
\
**Break up complex requests into multiple prompts**
If you have a complex task that requires several steps, you might want to break it into multiple prompts, and port the output of each step into the prompt for the next step. That way you can fine tune the results (and temperature) for each step.For example, to provide a list of doctors for a patient with a specific symptom, you might break it into these steps:
1. Given the patient’s symptoms, determine what the likely problem is.
2. Given the problem, decide which type of doctor handles that type of problem.
3. Look up the list of doctors with that specialty and provide contact information for each.
**Break long inputs into smaller segments**\
When asking the model to process or rewrite very long documents (above 10k tokens),\
don’t submit them in a single pass. The model is likely to compress, summarize, or truncate the content. Instead, explicitly instruct it to divide the source into smaller, logically coherent segments (around 400–800 words each), aligned with natural boundaries like headings or topic shifts. Process each segment in order, preserving its detail and length, then stitch the outputs back together with light editing for flow.
Rewrite the following 12,000-word document while keeping the same level of detail and length.
Divide the following document into coherent sections of about 400–800 words. For each section, rewrite it with the same level of detail and approximately the same length, keeping the order intact. After processing all sections, combine the rewritten segments into a single continuous text.
# Quantization
Source: https://docs.ai21.com/docs/quantization
Quantization reduces model memory usage by representing weights with lower precision. Learn how to use quantization techniques with Jamba models for efficient inference and training.
## Overview
Jamba models support several quantization techniques:
* **FP8 Quantization**: 8-bit floating point weights for reduced memory footprint and efficient deployment
* **ExpertsInt8**: Innovative quantization for MoE models in vLLM deployment
* **8-bit Quantization**: Using BitsAndBytesConfig for training and inference
## FP8 Quantization (vLLM)
These models leverage pre-quantized FP8 weights, significantly reducing storage requirements and memory footprint while not compromising output quality.
FP8 quantization requires Hopper architecture GPUs such as NVIDIA H100 and NVIDIA H200.
### Pre-quantized Model Weights
#### Prerequisites
```bash theme={"system"}
pip install vllm>=0.6.5,<=0.8.5.post1
```
#### Implementation
```python theme={"system"}
from vllm import LLM, SamplingParams
llm = LLM(
model="ai21labs/AI21-Jamba-Mini-1.7-FP8",
max_model_len=100*1024,
)
```
```python theme={"system"}
sampling_params = SamplingParams(
temperature=0.4,
top_p=1.0,
max_tokens=100
)
prompts = ["Explain the advantages of FP8 quantization:"]
outputs = llm.generate(prompts, sampling_params)
print(outputs[0].outputs[0].text)
```
Pre-quantized FP8 models require no additional quantization parameters since the weights are already quantized.
## ExpertsInt8 Quantization (vLLM)
[ExpertsInt8](https://github.com/vllm-project/vllm/pull/7415) is an innovative and efficient quantization technique developed specifically for Mixture of Experts (MoE) models deployed in vLLM, including Jamba models. This technique enables:
* **Jamba Mini 1.7**: Deploy on a single 80GB GPU
* **Jamba Large 1.7**: Deploy on a single node of 8x 80GB GPUs
### Prerequisites
```bash theme={"system"}
pip install vllm>=0.6.5,<=0.8.5.post1
```
### Implementation
```python theme={"system"}
from vllm import LLM
llm = LLM(
model="ai21labs/AI21-Jamba-Mini-1.7",
max_model_len=100*1024,
quantization="experts_int8" # Enable ExpertsInt8 quantization
)
```
```python theme={"system"}
from vllm import SamplingParams
sampling_params = SamplingParams(
temperature=0.4,
top_p=0.95,
max_tokens=100
)
# Generate text
prompts = ["Explain the benefits of model quantization:"]
outputs = llm.generate(prompts, sampling_params)
print(outputs[0].outputs[0].text)
```
With ExpertsInt8 quantization, you can fit prompts up to 100K tokens on a single 80GB A100 GPU with Jamba Mini.
## 8-bit Quantization (Hugging Face)
With 8-bit quantization using BitsAndBytesConfig, it is possible to fit up to 140K sequence length on a single 80GB GPU.
### Prerequisites
```bash theme={"system"}
pip install transformers torch bitsandbytes accelerate
```
### Implementation
```python theme={"system"}
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
quantization_config = BitsAndBytesConfig(
load_in_8bit=True, # Enable 8-bit quantization
llm_int8_skip_modules=["mamba"] # Exclude Mamba blocks to preserve quality
)
```
```python theme={"system"}
model = AutoModelForCausalLM.from_pretrained(
"ai21labs/AI21-Jamba-Mini-1.7",
torch_dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
quantization_config=quantization_config
)
# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained("ai21labs/AI21-Jamba-Mini-1.7")
```
```python theme={"system"}
messages = [
{"role": "system", "content": "You are a helpful AI assistant."},
{"role": "user", "content": "What are the advantages of 8-bit quantization?"}
]
input_ids = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors='pt'
).to(model.device)
with torch.no_grad():
outputs = model.generate(
input_ids,
max_new_tokens=200,
temperature=0.7,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
response = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(response)
```
To maintain model quality, we recommend excluding Mamba blocks from quantization using `llm_int8_skip_modules=["mamba"]`.
# Responsible Use
Source: https://docs.ai21.com/docs/responsible-use-1
AI21 Studio provides open access to state-of-the-art language models that can be used to power a large variety of useful applications. We believe it is important to ensure that this technology is used in a responsible way, while allowing developers the freedom they need to experiment rapidly and deploy solutions at scale.\
\
In order to use AI21 Studio, you are required to comply with our [Terms of Service](https://www.ai21.com/terms-policies/terms-of-use/) and with the following [Usage Guidelines](#usage-guidelines). Provided you comply with these requirements, you may use AI21 Studio to power applications with live users without any additional approval. We reserve the right to limit or suspend your access to AI21 Studio at any time where we believe these terms or guidelines are violated.
## Usage Guidelines
1. AI21 Studio must not be used for any of the following activities:
1. Illegal activities, such as hate speech, gambling, child pornography or violating intellectual property rights;
2. Harassment, victimization, intimidation, fraud or spam;
3. Creation or dissemination of misinformation, promotion of self-harm, glorification of violent events or incitement of violence.
2. Your application may present content generated by AI21 Studio directly to humans (e.g., chatbots, content generation tools, etc). **In this case**, you are required to ensure the following:
1. No content generated by AI21 Studio will be posted **automatically** (without human intervention) to any public website or platform where it may be viewed by an audience greater than **100 people**.
### Example
This means that you can use AI21 Studio to build a bot for your team’s 7-person Slack channel. In contrast, you are not allowed to build a Twitter bot, unless each tweet is checked by a human before it is posted. You can build a customer service bot that interacts with any number of customers, assuming it's chatting with each human separately in a 1:1 conversation.
2. In any case, the first human to view text generated by AI21 Studio must not be led to believe that it was written by a human.
### Example
If you’re building a copywriting tool for marketing professionals, your users must be informed that the text proposed to them is machine generated. They are then free to use and present it as their own, at their discretion. As another example, if you’re building a chatbot, it must be clear to your users that they are corresponding with a machine rather than a live human.
3. Language models such as those accessible via AI21 Studio may generate inappropriate, biased, offensive or otherwise harmful content (see our [technical paper](https://arxiv.org/abs/2408.12570) for an evaluation of bias in our models). If your application is used by more than **100 people per month**, you must provide a method for users to **report generated text as harmful**. You should monitor these reports and respond to them appropriately.
### Example
You can build a demo, launch a closed beta, etc. without any special requirements, as long as it is accessed by fewer than 100 users per month. Once you exceed 100 monthly users, you must implement a “flag as inappropriate” button or some similar functionality to collect feedback.
3. The prompt text for any completion request must contain at least **60 characters** of text (about 10 words) **not written by your users**. This text should be crafted by you to produce the desired functionality for the user.
4. Language models such as those accessible via AI21 Studio can generate content that is biased against particular groups of people. You may not use AI21 Studio to power **automated decision making** where individuals may be denied benefits, refused access to a service or otherwise have their wellbeing substantially harmed based on protected characteristics.
5. AI21 Studio must not be used to classify or profile people based on protected characteristics (like racial or ethnic background, religion, political views, or private health data).
# Safety Research
Source: https://docs.ai21.com/docs/safety-research-1
AI21 Labs is on a mission to make reading and writing AI-first experiences, with machines working alongside humans as thought partners, thereby promoting human welfare and prosperity. To deliver its promise, this technology must be deployed and used in a responsible and sustainable way, taking into consideration potential risks, including malicious use by bad actors, accidental misuse and broader societal harms. We take these risks extremely seriously and put measures in place to mitigate them.\
\
AI safety is an important challenge with a large surface area, which we believe can be addressed most effectively by **working together**. We invite anyone interested in conducting research or otherwise promoting AI safety to contact us at [safety@ai21.com](mailto:safety@ai21.com) and explore opportunities for collaboration. We encourage members of the community to contact us at the same address to report bad experiences, vulnerabilities and suspected misuse of our products or to voice any other safety-related concerns.
# SDK
Source: https://docs.ai21.com/docs/sdk
AI21 offers Python and TypeScript libraries that simplify the process of using our LLMs' API.
## Get an API key
Before you can start using the SDK, you'll need to obtain your API key from [AI21 Studio](https://studio.ai21.com/v2?tab=maestro).\
See the full guide here: [**Create an API Key**](https://docs.ai21.com/docs/create-api-key)**.**
### Installation
Install ai21 Python SDK with your favorite package manager.
```shell Shell theme={"system"}
pip install ai21
```
### Authentication
There are two ways to authenticate with the SDK:\
\
**Option 1: Pass the API key directly**
```python theme={"system"}
from ai21 import AI21Client
client = AI21Client(api_key="your_api_key")
```
**Option 2: Use an environment variabl**e
Set the variable in your terminal:
```shellscript theme={"system"}
export AI21_API_KEY="your_api_key"
```
Then in your Python code, you can initialize the client without explicitly passing the key:
```python theme={"system"}
client = AI21Client()
```
For details on how authentication works at the HTTP level, see the [API Refrence Autentication](https://docs.ai21.com/reference/authentication).
### Example usage
This example demonstrates how to initialize the AI21 Python SDK client and send a simple chat message using the `jamba-mini` model.
```python theme={"system"}
from ai21 import AI21Client
from ai21.models.chat import ChatMessage
client = AI21Client(
# defaults to os.environ.get('AI21_API_KEY')
api_key='your_api_key',
)
system = "You're a support engineer in a SaaS company"
messages = [
ChatMessage(content=system, role="system"),
ChatMessage(content="Hello, I need help with a signup process.", role="user"),
]
chat_completions = client.chat.completions.create(
messages=messages,
model="jamba-mini",
)
print(chat_completions.choices[0].message.content)
```
## Python SDK
For more examples and detailed usage, check out our Python library GitHub repository.
# Tokenization
Source: https://docs.ai21.com/docs/tokenization
Tokenization is the process of converting text into numerical tokens that language models can understand. Learn how to use AI21's tokenizer for Jamba models with practical examples.
## What is Tokenization?
Tokenization is both the first and final step in language model processing. Since machine learning models can only work with numerical data, text must be converted into numbers that models can understand and manipulate.
The tokenization process breaks down text into smaller units called tokens, which can represent:
* **Words or subwords**: "hello" → `[15496]`
* **Characters**: "AI" → `[32, 73]`
* **Byte-level representations**: For handling any Unicode text across all languages
Each token is assigned a unique numerical ID from the model's vocabulary.
## When is Tokenization Used?
Tokenization serves as both the entry point and exit point of text processing in language models. Since models can only work with numerical data, text must be converted into tokens with corresponding numerical indices from the tokenizer's vocabulary.
In a standard language model workflow:
1. **Encoding Phase**: We first convert input text into tokens using a tokenizer. Each token receives a unique index number that the model can process.
2. **Model Processing**: The tokenized input flows through the model architecture:
* **Embedding layer**: Transforms tokens into dense vector representations that capture semantic relationships
* **Transformer blocks**: Process these vectors to understand context, relationships, and generate meaningful responses
3. **Decoding Phase**: Finally, we convert the model's output tokens back into readable text by mapping token indices back to their corresponding words or subwords using the tokenizer's vocabulary.
This encode → process → decode cycle ensures seamless conversion between human language and machine-readable formats, enabling effective communication with language models.
## AI21's Tokenizer
We provides a [AI21-Tokenizer](https://github.com/AI21Labs/ai21-tokenizer) specifically engineered for Jamba models.
### Key Features
* **Jamba Mini and Large support**
* **Async/sync operations**:
* **Production-ready**: Enterprise-grade reliability
## Installation
### Prerequisites
To use tokenizers for Jamba, you'll need access to the relevant model's HuggingFace repository.
### Install the Tokenizer
```bash theme={"system"}
pip install ai21-tokenizer
```
## Model-Specific Tokenizers
Choose the appropriate tokenizer for your Jamba model:
```python theme={"system"}
from ai21_tokenizer import Tokenizer, PreTrainedTokenizers
tokenizer = Tokenizer.get_tokenizer(PreTrainedTokenizers.JAMBA_MINI_TOKENIZER)
text = "Jamba Mini model says hello"
encoded = tokenizer.encode(text)
print(f"Jamba Mini encoded: {encoded}")
```
```python theme={"system"}
from ai21_tokenizer import Tokenizer, PreTrainedTokenizers
tokenizer = Tokenizer.get_tokenizer(PreTrainedTokenizers.JAMBA_LARGE_TOKENIZER)
text = "Jamba Large model says hello"
encoded = tokenizer.encode(text)
print(f"Jamba Large encoded: {encoded}")
```
## Basic Usage
```python theme={"system"}
from ai21_tokenizer import Tokenizer
# Create tokenizer (defaults to Jamba Mini)
tokenizer = Tokenizer.get_tokenizer()
# Convert text to token IDs
text = "Hello, world!"
encoded = tokenizer.encode(text)
print(f"Encoded: {encoded}")
# Output: Encoded: [15496, 11, 1917, 0]
```
```python theme={"system"}
from ai21_tokenizer import Tokenizer
tokenizer = Tokenizer.get_tokenizer()
# Convert token IDs back to text
decoded = tokenizer.decode(encoded)
print(f"Decoded: {decoded}")
# Output: Decoded: Hello, world!
```
## Asynchronous Usage
For high-performance/server applications, use the async tokenizer:
```python theme={"system"}
import asyncio
from ai21_tokenizer import Tokenizer
async def main():
tokenizer = await Tokenizer.get_async_tokenizer()
text = "Async tokenization for async operations!"
encoded = await tokenizer.encode(text)
decoded = await tokenizer.decode(encoded)
print(f"Original: {text}")
print(f"Encoded: {encoded}")
print(f"Decoded: {decoded}")
asyncio.run(main())
```
## Practical Use Cases
* **Cost estimation**: Calculate API usage costs based on token consumption
* **Prompt optimization**: Ensure prompts fit within model context limits
For more advanced usage examples, visit the [AI21 tokenizer examples](https://github.com/AI21Labs/ai21-tokenizer/tree/main/examples) folder.
# Training Data
Source: https://docs.ai21.com/docs/training-data-1
The model training dataset is composed of text posted or uploaded to the internet. The internet data that it has been trained on includes a filtered and curated version of the CommonCrawl dataset, Wikipedia the BookCorpus dataset and arXiv/Stack Exchange.\
\
In the pretraining process, we excluded sites with robot files indicating the presence of copyright material and/or PII. All training data was filtered using our toxicity, safety and similarity detection technology and processes to improve the safety, accuracy and reliability of the output of the model.
The creation of a training dataset can be viewed as a pipeline consisting of selection, curation, filtering, augmentation and ingestion. This process is iterative and involves both human and machine evaluation in each phase of the pipeline. Employees of AI21 are involved in every phase and third-party organizations are used in the filtering and augmentation phases of the data pipeline and in later testing (e.g. red-teaming) to provide external review and validation. Due diligence has been performed on the business practices of these third-party organizations including locations, wages, working conditions and protections.
Customers of AI21 and government agencies can request additional details about these organizations, their operations and the roles they play and the instructions given to them. Considering the training data used by the model, it follows that its outputs can be interpreted as being representative of internet-connected populations. Most of the internet-connected population is from industrialized countries, wealthy, younger, and male, and is predominantly based in the United States. This means that less-connected, non-English-speaking and poorer populations are less represented in the outputs of the system.
Customers can compensate for some of these limitations by adding training data of their own, but the underlying language model inherently contains bias based on its pre-training data.
# Troubleshooting & Performance Optimization
Source: https://docs.ai21.com/docs/troubleshooting-performance
Resolve common issues and optimize performance for AI21's Jamba model deployments
## Overview
This guide helps you troubleshoot common deployment issues and optimize performance for AI21's Jamba models across different deployment scenarios.
Before troubleshooting, ensure you're using the recommended vLLM version `v0.6.5` to `v0.8.5.post1`.
## Troubleshooting
### Memory Issues
**Symptoms:**
* CUDA out of memory errors
* Process killed by system
**Solutions:**
```bash theme={"system"}
--quantization="experts_int8" # Reduces memory usage (recommended)
--tensor-parallel-size=8 # Number of GPUs to use (1-8)
--max-model-len=128000 # Reduce context length if needed (max 256K)
--gpu-memory-utilization=0.8 # Limits GPU memory usage
```
**Symptoms:**
* Inconsistent OOM errors
* Memory usage appears lower than expected
**Solutions:**
```bash theme={"system"}
--max-num-seqs=50 # Controls memory per request (increase/decrease to tune)
```
### Model Loading Issues
**Storage Recommendations:**
* Network Storage: >1 GB/s bandwidth
* RAM Disk: Load model from RAM if possible
### Performance Issues
**When to Use:**
* Long input sequences (>8K tokens)
* High memory pressure during prefill
* Mixed sequence lengths in batches
**Configuration:**
```bash theme={"system"}
--enable-chunked-prefill # Enables chunked prefill
--max-num-batched-tokens=8192 # Adjust based on GPU memory
```
## Getting Help
If you need support, please contact our team at **[support@ai21.com](mailto:support@ai21.com)** with the following information:
**Environment Details:**
* Hardware specifications (GPU model, memory, CPU)
* Software versions (vLLM, CUDA, drivers)
* Full vLLM command
**Diagnostics:**
* Full error messages and stack traces
* GPU utilization logs (`nvidia-smi` output)
# Pricing
Source: https://docs.ai21.com/docs/usage-cost
Usage costs depend on whether you’re accessing AI21 models through the AI21 platform (REST endpoints, the SDK, or the AI21 playground) or through a third-party host (a cloud platform), which endpoints you use, and whether you have a custom payment plan with AI21.
Pricing is calculated either by token:
* **Cost per token**, where input and output tokens may be priced differently. Token usage counts are given in the usage field in API responses.
## AI21 platform usage
Usage costs for the AI21 Platform (REST endpoints, the SDK, or the AI21 playground) are [given here](https://www.ai21.com/pricing).
### Free trial usage
New accounts are given a \$10 credit good for three months on the AI21 platform. You can use your credit for all usage of the APIs, the SDK, and the playground. Once your trial usage is exceeded or expires, you must provide valid billing information in your account to continue using the AI21 platform.
### Subsequent usage
Once you have exceeded your free trial period or credit, you must provide valid billing information to continue using AI21 services. If you have large usage expectations, consider signing up for a [customized plan](https://www.ai21.com/pricing).
Usage is charged monthly from account creation date. Per-token costs are per thousand (K) or per million (M), with no rounding. So 1,500 tokens at \$0.10/K tokens cost (1,500 / 1,000) \* \$0.10 = \$0.15
### See your AI21 Platform usage and billing
* **See your token counts and costs:** *Navigate to* **Account >[Model usage](https://studio.ai21.com/account/model-usage)**
* **See or change your billing plan:** *Navigate to* **Account >[Billing & Plans](https://studio.ai21.com/account/billing-plans)**
## Cloud usage
AI21 models are hosted on several cloud providers, including AWS SageMaker and Bedrock, and Microsoft Azure. Cloud usage costs are set by the cloud provider. For example, [SageMaker](https://aws.amazon.com/marketplace/pp/prodview-dkwy6chb63hk2) charges rates based on usage time and instance type, whereas [Bedrock](https://aws.amazon.com/bedrock/pricing/) charges a per-token fee.
***
# Use Cases
Source: https://docs.ai21.com/docs/use-cases-validated-output
Practical applications where the Validated Output delivers significant value.
## Data Extraction and Structured Output
\
Travel Booking Information Extraction
**Challenge**: Extract structured information from conversational data while handling missing information appropriately and avoiding hallucinations.
**Scenario**: Processing travel booking conversations to extract key booking details in a structured format.
**Performance Comparison: Chat Model**
We evaluate the performance of a state-of-the-art chat model in two configurations:
* Chat model alone (baseline)
* Chat model enhanced with AI21 Maestro
### Baseline Chat Model Output (Without AI21 Maestro)
Despite receiving explicit instructions, the chat model fails to properly handle partial information and makes unauthorized inferences:
**Input Prompt**
```
You are an expert in extracting structured information from unstructured text, specifically travel booking conversations.
You are given a conversation as plain text between a user and an AI travel agent. Extract and return the following fields:
Name: Full name of the traveler
HotelName: Name of the hotel selected by the traveler
CheckInDate: Check-in date in YYYY-MM-DD format if the date can fully be extracted
CheckOutDate: Check-out date in YYYY-MM-DD format if the date can fully be extracted
City: City where the hotel is located
Country: Country where the hotel is located
NumberOfGuests: Total number of guests
Requirements:
1. Return a bulleted list where each field is a bullet
2. If any field is not mentioned or can’t be fully extracted from the conversation, fill it with "NA
3. Do not invent, infer, or assume any details that are not explicitly stated in the conversation.
Conversation:
User: Hi there! I'm planning a trip to Rome next month and need help booking a hotel.
AI: I'd be happy to help! Could you tell me your travel dates and how many people will be staying?
User: Sure. I'll be arriving on July 10th and checking out on July 15th. It'll be just me and my wife, so 2 guests.
AI: Great. Would you like me to recommend some hotels in Rome, Italy?
User: Yes, please.
AI: Here are a few options:
Hotel Artemide
The Rome EDITION
Hotel Nazionale
Do any of these interest you?
User: Hotel Artemide sounds perfect. Let's go with that.
AI: Excellent choice! May I have your full name for the reservation?
User: It's Jonathan Rivera.
AI: Thank you, Jonathan. I've reserved a room at Hotel Artemide in Rome, Italy for 2 guests from July 10 to July 15.
```
**Chat Model Output**
```
- Name: Jonathan Rivera
- HotelName: Hotel Artemide
- CheckInDate: 2023-07-10 ❌ VIOLATION: Inferred year (2023) not in conversation
- CheckOutDate: 2023-07-15 ❌ VIOLATION: Inferred year (2023) not in conversation
- City: Rome
- Country: Italy
- NumberOfGuests: 2
```
### **Issues Identified**
* **Instruction Violation**: Added year "2023" despite explicit instruction not to infer
* **Compliance Failure**: Ignored "fill with NA" requirement for incomplete dates
### AI21 Maestro Performance
**Input Prompt**
```
You are an expert in extracting structured information from unstructured text, specifically travel booking conversations.
You are given a conversation as plain text between a user and an AI travel agent. Extract and return the following fields:
Name: Full name of the traveler
HotelName: Name of the hotel selected by the traveler
CheckInDate: Check-in date in YYYY-MM-DD format if the date can fully be extracted
CheckOutDate: Check-out date in YYYY-MM-DD format if the date can fully be extracted
City: City where the hotel is located
Country: Country where the hotel is located
NumberOfGuests: Total number of guests
Conversation:
User: Hi there! I'm planning a trip to Rome next month and need help booking a hotel.
AI: I'd be happy to help! Could you tell me your travel dates and how many people will be staying?
User: Sure. I'll be arriving on July 10th and checking out on July 15th. It'll be just me and my wife, so 2 guests.
AI: Great. Would you like me to recommend some hotels in Rome, Italy?
User: Yes, please.
AI: Here are a few options:
Hotel Artemide
The Rome EDITION
Hotel Nazionale
Do any of these interest you?
User: Hotel Artemide sounds perfect. Let's go with that.
AI: Excellent choice! May I have your full name for the reservation?
User: It's Jonathan Rivera.
AI: Thank you, Jonathan. I've reserved a room at Hotel Artemide in Rome, Italy for 2 guests from July 10 to July 15.
```
**Requirements**
1. If any field is not mentioned or can’t be fully extracted from the conversation, fill it with "NA”
2. Do not invent, infer, or assume any details that are not explicitly stated in the conversation.
3. return a bulleted list where each field is a bullet
**Output**
```
- Name: Jonathan Rivera
- HotelName: Hotel Artemide
- CheckInDate: NA ✅ CORRECT: Year not specified, properly marked as NA
- CheckOutDate: NA ✅ CORRECT: Year not specified, properly marked as NA
- City: Rome
- Country: Italy
- NumberOfGuests: 2
```
# HTTP tools
Source: https://docs.ai21.com/docs/user-provided-tools-http
Description of your new file.
## Overview
HTTP tools allow AI21 Maestro to call external services directly using standard HTTP requests. Unlike MCP, which supports dynamic discovery and standardized communication, HTTP tools are statically defined and managed by the user. They are best for simple or one-off integrations that don’t require a full MCP server.
## When to Use HTTP Tools
* When you want a **quick integration** without building an MCP Server.
* When you have standalone endpoints that don’t require cataloged discovery.
## Defining an HTTP Tool
Each HTTP tool must include a `function` definition and an `endpoint` configuration.
* The `function` field describes **what the tool does** and **what input it expects**.
| Field | Description |
| :--------------------- | :----------------------------------------- |
| `type` | Must be `"http"` |
| `function.name` | Unique function identifier (per execution) |
| `function.description` | What the tool does and when to use it |
| `function.parameters` | JSON schema for input parameters |
* The `endpoint` field defines how to call it over HTTP:
| Field | Description |
| :------------------------------ | :------------------------------------------ |
| `endpoint.url` | The HTTP URL that will be called via POST |
| `endpoint.headers` *(optional)* | Custom headers, typically for authorization |
**JSON Example of HTTP Tool definition:**
```json theme={"system"}
{
"type": "http",
"function": {
"name": "get_weather",
"description": "Get current temperature for a given location.",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City and country e.g. Bogotá, Colombia"
}
},
"required": [
"location"
],
"additionalProperties": false
}
},
"endpoint": {
"url": "https://my.endpoint.com",
"headers": {
"Authorization": "Bearer {api_key}"
}
}
}
```
## Request Flow & Authentication
Once defined, AI21 Maestro can call this tool by sending a `POST` request to the provided endpoint.
### **Example Request:**
Given that the user provided the example input above with the `get_weather` tool, and during an execution Maestro decides to call it with "Tel-Aviv" as the location, the following HTTP request will be sent to the user's endpoint:
URL: `https://my.endpoint.com/get_weather`
Method: `POST`
Headers: `Authorization: Bearer {api_key}`
Body:
```json theme={"system"}
{
"location": "Tel-Aviv"
}
```
### Authentication
Headers are defined statically in the `endpoint.headers` object. This is where you can pass:
* Bearer tokens
* API keys
* Session headers or custom headers
## Best Practices
* **Use specific tool names** that clearly reflect their action (`get_user_profile`, `get_invoice`).
* **Include descriptions for all parameters**, not just the tool itself — this improves tool selection by the LLM.
* **Avoid overly broad or generic tools**; precise and focused tools tend to perform better.
* **Limit additional properties** unless the endpoint accepts flexible payloads.
# MCP Overview
Source: https://docs.ai21.com/docs/user-provided-tools-mcp
The [**Model Context Protocol (MCP)**](https://modelcontextprotocol.io/docs/getting-started/intro) is an open protocol that standardizes how applications provide context to large language models (LLMs). By using MCP, you can reduce development time and complexity when building or integrating AI applications and agents.\
\
[**Remote MCP servers**](https://modelcontextprotocol.io/docs/develop/connect-remote-servers#understanding-remote-mcp-servers) extend AI applications’ capabilities beyond your local environment, giving them access to internet-hosted tools, services, and data sources. Connecting to these servers allows AI assistants to handle complex, multi-step projects with real-time access to external resources. The main benefit of remote MCP servers is their accessibility. Unlike local servers that must be installed and configured on individual devices, remote servers can be reached from any MCP client with an internet connection. Remote MCP servers allow AI21 Maestro to connect with internet-hosted tools, services, and data sources in real time. This makes them ideal for complex, multi-step workflows that require secure server-side processing and authentication. \
\
When you define an MCP tool in the `tools` parameter, AI21 Maestro acts as an MCP client. It authenticates with the server, retrieves the list of available tools defined there, and can invoke them during a session to fetch data or perform actions.
AI21 Maestro supports only remote MCP servers, not local ones.
## When to use MCP
Use MCP when you need:
* **Dynamic discovery** of available tools from a remote server.
* **Access to multiple related tools** (for example, a Jira MCP server exposing several Jira-related functions).
* **Secure, server-side integration** with authentication.
* **Reusable integrations** that can be shared across multiple sessions or projects.
## **How to Register an MCP Server**
### **Required fields:**
| Field | Description |
| :--------------------------- | :----------------------------------------------------------------- |
| `type` | Must be `"mcp"` |
| `server_label` | Unique name for the MCP server within the session |
| `server_url` | Fully qualified URL of the MCP server endpoint |
| `headers` *(optional)* | Authentication headers (e.g., bearer tokens) |
| `allowed_tools` *(optional)* | List of tool names to import; if omitted, all tools will be loaded |
### **JSON Example**
```json theme={"system"}
{
"type": "mcp",
"server_label": "Jira",
"server_url": "https://mcp.atlassian.com/v1/sse",
"headers": {
"Authorization": "Bearer "
},
"allowed_tools": [
"getJiraIssue",
"searchJiraIssuesUsingJql",
"getVisibleJiraProjects"
]
}
```
### **Authentication**
MCP supports common authentication schemes used in HTTP-based APIs. You can include credentials via headers when defining the server.
### Supported Methods:
* **Bearer tokens** in the `Authorization` header
* **API keys** in custom headers
* **OAuth tokens** or session-based headers
## Best Practices
* **Use descriptive tool names** to make them easy for agents to understand and select.
* **Limit tools** with `allowed_tools` if your server exposes many capabilities.
* **Write complete descriptions** for functions and parameters—these are used by the model to decide when and how to invoke them.
* **Secure your endpoints** using proper authentication and avoid exposing unnecessary tools.
## Next steps
Build an Employee Server with[ FastMCP](https://docs.ai21.com/docs/mcp-server-setup-guide).
# Overview
Source: https://docs.ai21.com/docs/user-provided-tools-overview
AI21 Maestro’s User-Provided Tool Calling feature enables integration with external systems via user-defined tools. Instead of relying solely on built-in capabilities, AI21 Maestro can incorporate user-defined tools and invoke them when relevant to a task or context.
These tools allow AI21 Maestro to interact with external services such as internal APIs, third-party platforms, or custom business logic by executing HTTP requests or other defined operations. This extends its ability to retrieve information, perform actions, or delegate parts of the task to external systems. Each tool must include a description specifying its name, purpose, input parameters, and expected output. Accurate and complete metadata improves the model’s ability to identify when a tool should be used.
AI21 Maestro is responsible for calling the tools directly, without returning control to the user.\
There are two supported types of user-provided tools:
**HTTP Tool**
This tool type is essentially identical to the widespread “Function tool.” \
The user is responsible for passing the function’s name, description, and parameter descriptions. In addition, the user needs to provide an endpoint which AI21 Maestro will call with the relevant parameters. Find more details on how AI21 Maestro interacts with the endpoint here.
**MCP Tool**
Tools are hosted on a remote server that implements the Model Context Protocol (MCP). AI21 Maestro connects to this server to fetch the available tools and invokes them directly when needed. This model can optionally include approval steps before execution.\
Find more details about the MCP tool here.
## Limitations
Due to AI21 Maestro’s multi-path execution and planning architecture, it may explore more than one way to accomplish a task, which can result in multiple calls to the same tool.\
Because of this behavior:
**Action / Write tools are strongly discouraged.**\
If such tools are invoked in more than one path, it can lead to unintended duplication of actions, inconsistent states, or conflicting side effects.
**Retrieval / Read tools are recommended instead.**\
They are safer under multi-path exploration because they avoid the risks associated with side effects and state changes: they provide information without altering external systems or resources.
## Key Benefits
* **Enhanced Context:** AI21 Maestro can retrieve data directly from your systems to better answer your queries.
* **Real-Time Data**: AI21 Maestro can access your systems directly without needing to index the data in advance.
* **Support for Existing Infrastructure**: You can reuse your existing tools and services instead of transforming internal knowledge into ingestible documents for file search.
## Example: Tool Definition
Here’s a simplified **Python snippe**t of a tool definition for accessing JIRA issue data:
```python theme={"system"}
@jira_mcp.tool(tags={"jira", "read"})
async def get_issue(
ctx: Context,
issue_key: Annotated[str, Field(description="Jira issue key (e.g., 'PROJ-123')")],
...
) -> str:
"""Get details of a specific Jira issue including its Epic links and relationship information."""
```
This tool could be registered for either HTTP Tool or MCP usage, depending on where and how it is hosted. Note that this tool implicitly contains the following definitions:
1. Tool name
2. Tool description
3. Parameter names
4. Parameter types
5. Parameter descriptions
Every one of these definitions helps AI21 Maestro decide when to call the tool and with which parameters to do so.\
\
**Example: Tool definition in JSON**
```json theme={"system"}
{
"name": "get_issue",
"description": "Get details of a specific Jira issue including its Epic links and relationship information.",
"tags": ["jira", "read"],
"input_schema": {
"type": "object",
"properties": {
"issue_key": {
"type": "string",
"description": "Jira issue key (e.g., 'PROJ-123')"
}
},
"required": ["issue_key"]
},
"output_schema": {
"type": "string"
}
}
```
## Practical Considerations
* Keep tool descriptions complete and accurate.
* Limit the number of tools exposed per use case to improve selection accuracy.
* Ensure tool outputs are structured in a machine-readable format (e.g., JSON) and consistently shaped.
## MCP vs. HTTP: Choosing the Right Tool
| | **MCP** | **HTTP** |
| ----------------------------------------------------- | ---------------------------------------- | --------------------------------- |
| **Multiple related tools** (e.g., Jira, HubSpot) | Yes | No |
| **Providing tool definitions** | The MCP provides them via tool discovery | The user provides the description |
| **The user provides the description** | No | Yes |
# Data Connectors
Source: https://docs.ai21.com/docs/using-data-connectors
How to Use Data Connectors in AI21 Studio
## Overview
Data Connectors allow you to ingest external data into your AI21 Maestro workspace. Once connected, this data can then be indexed and used to ground responses in external knowledge, enabling your AI applications to provide more accurate, relevant, and up-to-date answers.
Through the AI21 Studio UI, users can easily connect data sources manage ingestion, and monitor indexing via the **File Library**.
This guide walks you through the process of using Data Connectors via the AI21 Studio.
## Available Data Connectors for Enterprise Users
Below is a complete list of all data connectors available to enterprise users.
* Amazon S3
* Box
* Confluence
* Dropbox
* Google Drive
* Browse Website
* Intercom
* OneDrive
* Microsoft Teams
* Notion
* ServiceNow
* SharePoint
* Zendesk
## How to Use Data Connectors: Step by Step
### Step 1: Open the Data Connectors Section
Navigate to [Data Connectors](https://studio.ai21.com/v2/workspaces/429d8ce9-0f35-4626-a48c-231d28db34a8/integrations) from the left-hand menu in AI21 Studio.
### Step 2: Choose the Data Connector
On the Data Connectors page, you'll see a list of available integrations.
* **Browse Website** is available for all users.
* **Enterprise customer**s are eligible for all connectors.
* If you require any additional connectors, please contact our [support team](mailto:support@ai21.com)
### **Configure File Storage Data Connectors**
1. Toggle on the requested **File Storage** connector.
2. Authorize the requested data connector.
3. Select and authenticate your user account.
### **Configure Browse Website Data Connector**
You have two options to connect the Index Website:
**Option A: Sitemap XML**
Paste a sitemap link to ingest content automatically.
Only the first 100 links in the sitemap will be indexed.
**Option B: Specific URLs** \
Paste up to 100 URLs. \
When you past the urls, make sure to press enter.
**Advanced Settings** \
There are additional ingestion options, such as:
* **Block URL Types**
* **Define URL Types**
* **Specify Localization**
* **Content Load Delay**
* **Extract Files (e.g., PDFs, Docs)**
These advanced features are available to Enterprise plan users. [Contact Sales](https://www.ai21.com/contact-sales/) for access.
### **Step 3: Click ‘Connect’**
Once you’ve pasted a sitemap or the URLs, click **Connect** to start ingesting the data.
### **Step 4: View Your Connected Data**
Once the data is connected you’ll see a checkmark (✓) next to it. Navigate to the **File Library** to see your connected data. It will appear in a table showing source, type, and ingestion status.
## Next Steps
Use [Data Connectors](https://studio.ai21.com/v2?tab=data_connectors) in AI21 Stuido.
# vLLM
Source: https://docs.ai21.com/docs/vllm
Deploy AI21's Jamba models using vLLM in your own environment. vLLM is an open-source library for high-throughput LLM inference and serving.
## Overview
This guide walks you through self-deploying AI21's Jamba models in your own infrastructure using [vLLM](https://github.com/vllm-project/vllm). Choose the deployment method that best fits your needs.
We recommend using vLLM version `v0.6.5` to `v0.8.5.post1` for optimal performance and compatibility.
## Prerequisites
For detailed information about hardware support and GPU requirements, see the [vLLM GPU Installation Guide](https://docs.vllm.ai/en/latest/getting_started/installation/gpu.html).
### System Requirements
* **Model Size:** 97GB
* **Compute Capability** 7.5+
* **Model Size:** 743GB
* **Compute Capability** 7.5+
* **Model Size:** 55GB
* **Compute Capability** 7.5+
* **Model Size:** 560GB
* **Compute Capability** 7.5+
* **Model Size:** 52GB
* **Compute Capability** 9.0+
* **Model Size:** 396GB
* **Compute Capability** 9.0+
## Deployment Options
### Option 1: vLLM Direct Usage
Create a Python virtual environment and install the vLLM package (version `≥0.6.5, ≤0.8.5.post1` to ensure maximum compatibility with all Jamba models).
```bash theme={"system"}
# Create and activate virtual environment
python -m venv vllm-env
source vllm-env/bin/activate
# Install vLLM
pip install vllm>=0.6.5,<=0.8.5.post1
```
Authenticate on the HuggingFace Hub using your access token `$HF_TOKEN`:
```bash theme={"system"}
huggingface-cli login --token $HF_TOKEN
```
Launch vLLM server for API-based inference:
**Start the vLLM server:**
```bash theme={"system"}
vllm serve ai21labs/AI21-Jamba-Mini-1.7 \
--quantization="experts_int8" \
--enable-auto-tool-choice \
--tool-call-parser jamba
```
**Test the API:**
```bash cURL theme={"system"}
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
-d '{
"model": "ai21labs/AI21-Jamba-Mini-1.7",
"messages": [
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
]
}'
```
```python Python theme={"system"}
import httpx
url = "http://localhost:8000/v1/chat/completions"
headers = {
"Content-Type": "application/json"
}
data = {
"model": "ai21labs/AI21-Jamba-Mini-1.7",
"messages": [
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
],
}
response = httpx.post(url, headers=headers, json=data)
print(response.json())
```
```python AI21 Python SDK theme={"system"}
from ai21 import AI21Client
client = AI21Client(
api_key="dummy-key", # Not needed for local vLLM server
api_host="http://localhost:8000/v1"
)
response = client.chat.completions.create(
model="ai21labs/AI21-Jamba-Mini-1.7",
messages=[
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
]
)
print(response.choices[0].message.content)
```
**Start the vLLM server:**
```bash theme={"system"}
vllm serve ai21labs/AI21-Jamba-Large-1.7 \
--quantization="experts_int8" \
--tensor-parallel-size=8 \
--enable-auto-tool-choice \
--tool-call-parser jamba
```
**Test the API:**
```bash cURL theme={"system"}
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
-d '{
"model": "ai21labs/AI21-Jamba-Large-1.7",
"messages": [
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
]
}'
```
```python Python theme={"system"}
import httpx
url = "http://localhost:8000/v1/chat/completions"
headers = {
"Content-Type": "application/json"
}
data = {
"model": "ai21labs/AI21-Jamba-Large-1.7",
"messages": [
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
],
}
response = httpx.post(url, headers=headers, json=data)
print(response.json())
```
```python AI21 Python SDK theme={"system"}
from ai21 import AI21Client
client = AI21Client(
api_key="dummy-key", # Not needed for local vLLM server
api_host="http://localhost:8000/v1"
)
response = client.chat.completions.create(
model="ai21labs/AI21-Jamba-Large-1.7",
messages=[
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
]
)
print(response.choices[0].message.content)
```
**Start the vLLM server:**
```bash theme={"system"}
vllm serve ai21labs/AI21-Jamba-Mini-1.7-FP8 \
--enable-auto-tool-choice \
--tool-call-parser jamba
```
**Test the API:**
```bash cURL theme={"system"}
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
-d '{
"model": "ai21labs/AI21-Jamba-Mini-1.7-FP8",
"messages": [
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
]
}'
```
```python Python theme={"system"}
import httpx
url = "http://localhost:8000/v1/chat/completions"
headers = {
"Content-Type": "application/json"
}
data = {
"model": "ai21labs/AI21-Jamba-Mini-1.7-FP8",
"messages": [
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
],
}
response = httpx.post(url, headers=headers, json=data)
print(response.json())
```
```python AI21 Python SDK theme={"system"}
from ai21 import AI21Client
client = AI21Client(
api_key="dummy-key", # Not needed for local vLLM server
api_host="http://localhost:8000/v1"
)
response = client.chat.completions.create(
model="ai21labs/AI21-Jamba-Mini-1.7-FP8",
messages=[
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
]
)
print(response.choices[0].message.content)
```
**Start the vLLM server:**
```bash theme={"system"}
vllm serve ai21labs/AI21-Jamba-Large-1.7-FP8 \
--tensor-parallel-size=8 \
--enable-auto-tool-choice \
--tool-call-parser jamba
```
**Test the API:**
```bash cURL theme={"system"}
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
-d '{
"model": "ai21labs/AI21-Jamba-Large-1.7-FP8",
"messages": [
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
]
}'
```
```python Python theme={"system"}
import httpx
url = "http://localhost:8000/v1/chat/completions"
headers = {
"Content-Type": "application/json"
}
data = {
"model": "ai21labs/AI21-Jamba-Large-1.7-FP8",
"messages": [
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
],
}
response = httpx.post(url, headers=headers, json=data)
print(response.json())
```
```python AI21 Python SDK theme={"system"}
from ai21 import AI21Client
client = AI21Client(
api_key="dummy-key", # Not needed for local vLLM server
api_host="http://localhost:8000/v1"
)
response = client.chat.completions.create(
model="ai21labs/AI21-Jamba-Large-1.7-FP8",
messages=[
{"role": "user", "content": "Who was the smartest person in history? Give reasons."}
]
)
print(response.choices[0].message.content)
```
In offline mode, vLLM loads the model to perform batch inference tasks in a one-off, standalone manner.
```python Jamba Mini theme={"system"}
from vllm import LLM
from vllm.sampling_params import SamplingParams
model_name = "ai21labs/AI21-Jamba-Mini-1.7"
sampling_params = SamplingParams(max_tokens=1024)
llm = LLM(
model=model_name,
quantization="experts_int8",
)
messages = [
{
"role": "user",
"content": "Who was the smartest person in history? Give reasons.",
}
]
res = llm.chat(messages=messages, sampling_params=sampling_params)
print(res[0].outputs[0].text)
```
```python Jamba Large theme={"system"}
from vllm import LLM
from vllm.sampling_params import SamplingParams
model_name = "ai21labs/AI21-Jamba-Large-1.7"
sampling_params = SamplingParams(max_tokens=1024)
llm = LLM(
model=model_name,
quantization="experts_int8",
tensor_parallel_size=8,
)
messages = [
{
"role": "user",
"content": "Who was the smartest person in history? Give reasons."
}
]
res = llm.chat(messages=messages, sampling_params=sampling_params)
print(res[0].outputs[0].text)
```
```python Jamba Mini FP8 theme={"system"}
from vllm import LLM
from vllm.sampling_params import SamplingParams
model_name = "ai21labs/AI21-Jamba-Mini-1.7-FP8"
sampling_params = SamplingParams(max_tokens=1024)
llm = LLM(
model=model_name,
)
messages = [
{
"role": "user",
"content": "Who was the smartest person in history? Give reasons.",
}
]
res = llm.chat(messages=messages, sampling_params=sampling_params)
print(res[0].outputs[0].text)
```
```python Jamba Large FP8 theme={"system"}
from vllm import LLM
from vllm.sampling_params import SamplingParams
model_name = "ai21labs/AI21-Jamba-Large-1.7-FP8"
sampling_params = SamplingParams(max_tokens=1024)
llm = LLM(
model=model_name,
tensor_parallel_size=8,
)
messages = [
{
"role": "user",
"content": "Who was the smartest person in history? Give reasons."
}
]
res = llm.chat(messages=messages, sampling_params=sampling_params)
print(res[0].outputs[0].text)
```
### Option 2: Quick Start with Docker
For containerized deployment, use vLLM's official Docker image to run an inference server (refer to the [vLLM Docker documentation](https://docs.vllm.ai/en/latest/deployment/docker.html) for comprehensive details).
```bash theme={"system"}
docker pull vllm/vllm-openai:v0.8.5.post1
```
Launch vLLM in server mode with your chosen model:
```bash theme={"system"}
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HUGGING_FACE_HUB_TOKEN=" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:v0.8.5.post1 \
--model ai21labs/AI21-Jamba-Mini-1.7 \
--quantization="experts_int8" \
--enable-auto-tool-choice \
--tool-call-parser jamba
```
```bash theme={"system"}
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HUGGING_FACE_HUB_TOKEN=" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:v0.8.5.post1 \
--model ai21labs/AI21-Jamba-Large-1.7 \
--quantization="experts_int8" \
--tensor-parallel-size=8 \
--enable-auto-tool-choice \
--tool-call-parser jamba
```
```bash theme={"system"}
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HUGGING_FACE_HUB_TOKEN=" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:v0.8.5.post1 \
--model ai21labs/AI21-Jamba-Mini-1.7-FP8 \
--enable-auto-tool-choice \
--tool-call-parser jamba
```
```bash theme={"system"}
docker run --runtime nvidia --gpus all \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HUGGING_FACE_HUB_TOKEN=" \
-p 8000:8000 \
--ipc=host \
vllm/vllm-openai:v0.8.5.post1 \
--model ai21labs/AI21-Jamba-Large-1.7-FP8 \
--tensor-parallel-size=8 \
--enable-auto-tool-choice \
--tool-call-parser jamba
```
Once the container is up and in healthy state, you will be able to test your inference using the same code samples as in the [Online Inference (Server Mode)](#online-inference-server-mode) section. Make sure to use the correct model identifier based on your chosen quantization approach.
If you prefer to use your own storage for model weights, you can download them from your self-hosted storage (e.g., AWS S3, Google Cloud Storage) and mount the local path to the container using `-v /path/to/model:/mnt/model/` and `--model="/mnt/model/"` instead of the HuggingFace model identifier.
## Next Steps
Deploy on AWS, Google Cloud, or Azure for production workloads
Optimize performance and resolve common deployment issues
Learn about the complete API interface and parameters
## Resources
* [AI21 Jamba Mini on HuggingFace](https://huggingface.co/ai21labs/AI21-Jamba-Mini-1.7)
* [AI21 Jamba Large on HuggingFace](https://huggingface.co/ai21labs/AI21-Jamba-Large-1.7)
* [vLLM Documentation](https://docs.vllm.ai/en/latest/index.html)
* [ExpertsInt8 Quantization Details](https://github.com/vllm-project/vllm/pull/7415)
# Walkthrough Guide
Source: https://docs.ai21.com/docs/walkthrough-guide
A comprehensive guide to AI21 Maestro’s Validated Output capabilities.
AI21 Maestro is an intelligent agentic system designed to handle complex AI workflows.\
This guide focuses specifically on the **Validated Output**, providing practical examples that range from basic usage to advanced scenarios.
## Understanding the Problem
Traditional LLM interactions often look like this:
```python Python theme={"system"}
# Traditional approach - unreliable
response = client.complete(
prompt="""Write a Python function that:
- Calculates fibonacci numbers
- Is under 10 lines
- Has proper docstrings
- Uses descriptive variable names"""
)
# Sometimes works perfectly, sometimes doesn't follow all constraints
```
**Common issues:**
* LLMs often fail to consistently meet all the individual requirements outlined in the promp
* There is no visibility into *which* requirements were not met
* Requires manual trial and error to achieve the desired output
## How AI21 Maestro Works
Maestro's instruction following enhancer uses a **Generate → Validate → Fix** cycle:
1. **Generate**: Creates initial response following your requirements
2. **Validate**: Evaluates and scores each requirement (0.0 to 1.0)
3. **Fix**: Refines output for requirements that scored \< 1.0
4. **Repeat**: Continues until all requirements are met or budget is exhausted
This systematic approach to instruction following is part of Maestro's broader agentic architecture, designed to handle complex workflows with reliability and precision.
```text text theme={"system"}
Input + Requirements → Generate → Validate → Fix → Final Output + Report
↑ ↓
← ← ← ← ←
```
## Using the API
**The Input parameter**
You can pass a string to Maestro as an input and it will be treated as a user message.
```python Python theme={"system"}
from ai21 import AI21Client
client = AI21Client(api_key="your-api-key")
# The following function will block until the default timeout is reached
client.beta.maestro.runs.create_and_poll(
input="Explain quantum computing to a 10-year-old",
requirements=[
{
"name": "reading_level",
"description": "Use simple words appropriate for a 10-year-old"
},
{
"name": "length",
"description": "Keep explanation under 100 words"
}
]
)
```
Alternatively you can pass an input as an array of message to support multiple turns in a conversation.
```python Python theme={"system"}
input=[
{
"role": "user",
"content": "Explain quantum computing to a 10-year-old",
},
{
"role": "assistant",
"content": 'Quantum computing is like a super-smart computer that uses tiny things called "qubits" instead of regular bits. While regular bits are like tiny switches that can be off (0) or on (1), qubits can be both at the same time! This helps quantum computers solve really hard problems much faster than normal computers by trying many possibilities at once',
},
{
"role": "user",
"content": "Translate this to spanish",
},
],
```
### System Prompt
The `system_prompt` defines the **agent’s identity, operating principles, and boundaries** before it processes any input.\
It guides how Maestro **interprets inputs, chooses tools, and reasons** throughout the run.
**Use it to:**
* Define the agent’s **role and identity**\
*Example:* "You are a cautious financial journalist. Verify all data before reporting"
* Provide **context or environment**\
*Example:* Today’s date is November 10, 2025. User location: New York.
* Define **behavioral rules**\
*Example:* "Always verify numbers from reliable sources before reporting. If data is unclear, ask a clarifying question."
```python theme={"system"}
run = client.beta.maestro.runs.create_and_poll(
system_prompt="You are a cautious financial journalist. Verify all data before reporting.",
input="Write a brief update on today's top stock movements.",
requirements=[
{"name": "word_limit", "description": "No more than 120 words"},
{"name": "tone", "description": "Neutral, professional tone"}
],
budget="medium",
tools=[
{
"type": "web_search",
"urls": ["https://finance.yahoo.com", "https://www.reuters.com"]
}
],
include=["requirements_result"]
)
print(run.result)
print(run.requirements_result)
```
## Working with Requirements
### Writing Effective Requirements
**Good Requirements:**
```python Python theme={"system"}
requirements = [
{
"name": "word_count",
"description": "Response must be exactly between 150-200 words"
},
{
"name": "json_format",
"description": "Output must be valid JSON with 'title' and 'content' fields"
},
{
"name": "no_technical_jargon",
"description": "Avoid technical terms; explain concepts in plain English"
}
]
```
**Requirements to Avoid:**
```python Python theme={"system"}
# Too vague
{"name": "good_quality", "description": "Make it good"}
# Contradictory
{"name": "short_and_detailed", "description": "Be brief but very detailed"}
# Unmeasurable
{"name": "creative", "description": "Be creative and original"}
```
### Requirement Categories
**Format Requirements:**
```python Python theme={"system"}
{
"name": "markdown_format",
"description": "Use proper markdown with headers, bullet points, and code blocks"
}
```
**Content Requirements:**
```python Python theme={"system"}
{
"name": "include_examples",
"description": "Provide at least 2 concrete examples for each concept"
}
```
**Style Requirements:**
```python Python theme={"system"}
{
"name": "professional_tone",
"description": "Use formal business language, avoid contractions and slang"
}
```
**Technical Requirements:**
```python Python theme={"system"}
{
"name": "python_best_practices",
"description": "Follow PEP 8 style guidelines and use type hints"
}
```
## Requirements Report
Enable detailed reporting by including requirements\_result:
```python Python theme={"system"}
run = client.beta.maestro.runs.create_and_poll(
input="Write a product review for a smartphone",
requirements=[
{"name": "word_count", "description": "use 200-250 words"},
{"name": "pros_and_cons", "description": "Include both pros and cons sections"},
{"name": "rating", "description": "End with a 1-5 star rating. For example: (★★★★☆)"}
],
include=["requirements_result"],
budget="low"
)
print(f"Result: {run.result}")
# Analyze the results
print(f"Overall Score: {run.requirements_result["score"]}")
print(f"Completion Reason: {run.requirements_result["finish_reason"]}")
print("Requirements Results:")
for req in run.requirements_result["requirements"]:
print(f" {req["name"]}: {req["score"]}")
print(f" Issue: {req["reason"]}")
```
**Sample Output Analysis**
```python Python theme={"system"}
# Example output
Overall Score: 0.67
Completion Reason: Budget exhausted
word_count: 1.0
pros_and_cons: 1.0
rating: 0.6
Issue: Rating format is '4 out of 5' instead of star format (★★★★☆)
```
This tells you:
* 2 out of 3 requirements were perfectly met
* The rating requirement needs refinement
* You might need a higher budget or clearer requirement
## Budget Control and Performance
Use the `budget` parameter to control how much computational effort AI21 Maestro applies when executing your task. Higher budgets improve reasoning reliability but increase latency and cost.
The snippet below shows how to set different budget levels in your Maestro run. Replace `task` and `requirements` with your own input values and make sure the `client` is initialized with your API key as shown in the [Quickstar](https://docs.ai21.com/docs/instruction-following-module#basic-usage).
**Budget Levels Explained**
```python Python theme={"system"}
# High Budget - Maximum reliability (~100 seconds for complex tasks)
run = client.beta.maestro.runs.create_and_poll(
input=task,
requirements=requirements,
budget="high"
)
# Medium Budget - Balanced approach (~60 seconds)
run = client.beta.maestro.runs.create_and_poll(
input=task,
requirements=requirements,
budget="medium"
)
# Low Budget - enhanced reliability but favors latency (~20 seconds)
run = client.beta.maestro.runs.create_and_poll(
input=task,
requirements=requirements,
budget="low"
)
```
## Using Third-Party Models
You can run Maestro tasks with both AI21 and third-party models.\
Use the `models` parameter to specify which model to run your task with.\
If no model is specified, Maestro will automatically select a suitable model based on the task requirements.
```python Python theme={"system"}
run = client.beta.maestro.runs.create_and_poll(
input=task,
requirements=requirements,
models=["gpt-4o"], # Specify preferred model
budget="high"
)
```
## Setting the Response Language
You can control the output language of Maestro’s response using the `response_language` parameter. For example, to receive the result in Spanish:
```python theme={"system"}
run = client.beta.maestro.runs.create_and_poll(
input=task,
requirements=requirements,
models=["jamba-mini"],
budget="medium",
response_language="spanish"
)
print(run.output_text)
```
# Web Search
Source: https://docs.ai21.com/docs/web-search-studio-guide
How to Use Web Search in AI21 Studio
Follow these steps to search the web for real-tine data and updates.
### Step 1: **Go to the Playground Section**
Navigate to Maestro [**Playground**](https://studio.ai21.com/v2?tab=maestro) from the left-hand menu in AI21 Studio.
### Step 2: Add a Tool
In the **Configuration** panel, go to the **Tools** section and click **Add**.
### **Step 3: Select Web Search**
In the tools dropdown, select **Web search**.
### Step 4: Specify the Websites You Want to Search
Enter the websites you want to use (for example: [*bbc.com*](http://bbc.com), [*techcrunch.com*](http://techcrunch.com)).
When no websites are specified, all websites will be searched.
### Save the Web Search
After verifying that the websites are correct, click **Save**.
# AI21 Labs Documentation
Source: https://docs.ai21.com/home
Welcome to AI21 Labs
Start building your AI solution with comprehensive guides and API references for AI21 Maestro and our Jamba Family of Open Models.
Terms and Policies
Publications
# Rate Limits
Source: https://docs.ai21.com/reference/api-rate-limits
Rate limits serve as essential mechanisms to control the flow of requests made to a system, ensuring its stability and fair usage. Rate limits restrict the number of requests that can be made within a specific time frame. Rate limits are commonly measured by **RPM** (requests per minute) or **RPS** (requests per second), indicating the maximum allowed number of requests a user or application can make during that time.
Below are the rate limits for all of our models when accessed on the AI21 Platform via our SDK or REST endpoints. If you require a higher limit for your particular use case, please don't hesitate to get in touch with us at [sales@ai21.com](mailto:sales@ai21.com).
### Cloud-based services
Cloud providers have their own rate limits. For instance, Amazon SageMaker rate limits are determined by the instance that you deploy to hold the model. Amazon Bedrock has their own pricing and rate limits for AI21 model usage. See your cloud provider's documentation for details.
## Foundation models
Foundation models have usage limits per second (RPS) and per minute (RPM):
| Foundation Model | RPS | RPM |
| ---------------- | --- | --- |
| Jamba Large | 10 | 200 |
| Jamba Mini | 10 | 200 |
***
# Authentication
Source: https://docs.ai21.com/reference/authentication
All API requests must include an `Authorization` header \
with a **Bearer token** for authentication.
To authenticate:
1. Use your API key as the token.
2. Include it in the request header as shown below:
**Header format:**\
Authorization: Bearer \
Replace `` with your actual key.\
\
See the [Create an API Key guide](https://docs.ai21.com/docs/create-api-key) for step-by-step instructions.
# Introduction
Source: https://docs.ai21.com/reference/introduction
The AI21 API reference provides the technical details you need to authenticate, send requests, and interpret responses when working with AI21’s APIs.\
\
It covers multiple products such as our: Jamba Family of Open Models, AI21 Maestro, which is AI21 Maestro is a dynamic planning system that determines the optimal sequence of actions to solve a given task during inference time.\
\
To get started, visit the [Authentication](https://docs.ai21.com/reference/authentication) page and review the endpoint-specific reference sections in the navigation menu.\
\
If you have any questions, feel free to reach out to our support team via [**email**](mailto:support@ai21.com) or click the chat icon in the lower right corner.
# Chat request
Source: https://docs.ai21.com/reference/jamba-1-6-api-ref
POST https://api.ai21.com/studio/v1/chat/completions
## Overview
The Jamba API provides access to a set of instruction-following chat models. It describes the details for interacting with the chat model via the API endpoint and provides specifications for request and response structures.
***
## Request body
The name of the model to use.\
You can call our model without specifying a version by using the following model names:
* `jamba-large`
* `jamba-mini`
For more information on the available model versions, [click here](/docs/jamba-foundation-models).
A list of messages representing the conversation history. The structure of the message object depends on the type:
An initial system message is optional but recommended to set the tone of the chat.
The role of the entity that is creating the message.
The content of the message.
Input provided by the user.
The role of the entity that is creating the message, in this case the user.
The content of the message.
Response generated by the model. Include this in your request to provide context for future answers.
The role of the entity that is creating the message, in this case the assistant.
The content of the message.
The function calls generated by the model, such as tool invocations.
The id of the tool call.
The type of tool.
The function invoked by the model.
The name of the function.
The parameters of the function as a JSON schema.
Contains the output of a tool. Add the function output for user-implemented tools to enable a user-friendly model response. If included, ensure an assistant message with a tool\_calls entry with a matching id exists.
The role of the entity that is creating the message, in this case the tool.
The content of the message.
The message is a response to this tool call.
A list of tools that the model can use when generating a response.\
Currently, the only function type tools are supported.
The type of tool. Currently, the only supported value is "function".
Describes a function to call. Currently, all functions must be described by the user; there are no built-in functions. An example function template is given below.
The name of the function.
Provide a complete description of what the function does, what it returns, and any limitations.
Each function parameter has a name, a type ("string", "integer", "float", "array", "boolean", or "enum"), and a description.
The document parameter accepts a list of objects, each containing multiple fields.
The content of this "document".
Key-value pairs describing the document:
Type of metadata, like 'author', 'date', 'url', etc. Should be things the model understands.
Value of the metadata.
An object defining the output format required from the model.\
Setting it to `{ "type": "json_object" }` activates JSON mode, ensuring the generated message adheres to valid JSON structure.
The maximum number of tokens the model can generate in its response.\
For Jamba models, the maximum allowed value is 4096 tokens.
Controls the variety of responses provided—a higher value results in more diverse answers.\
**Default:** 0.4, **Range:** 0.0–2.0\
[More information](https://ai21-demo.mintlify.app/v4.2/docs/large-language-models)
Limit the pool of next tokens in each step to the top N percentile of possible tokens, where 1.0 means the pool of all possible tokens, and 0.01 means the pool of only the most likely next tokens. \
**Default**: 1.0, **Range**: 0 \<= value \<=1.0\
[More information](/v4.2/docs/large-language-models)
End the message when the model generates one of these strings. The stop sequence is not included in the generated message. Each sequence can be up to 64K long, and can contain newlines as `n` characters.
* **Single stop string with a word and a period**: "monkeys."
* **Multiple stop strings and a newline**: \["cat", "dog", " .", "####", "\n"]
How many chat responses to generate. *Default:1, Range: 1 – 16.* \
**Notes**:
* If `n > 1`, setting `temperature = 0` will fail because all answers are guaranteed to be duplicates.
* `n` must be 1 when `stream = True`
Stream results one token at a time using server-sent events. Useful for long results to avoid long wait times. If `True`, `n` must be 1. Must be `False` if using `tools`.
```python Python theme={"system"}
from ai21 import AI21Client
from ai21.models.chat import ChatMessage
messages = [
ChatMessage(role="user", content="Hello how are you?"),
]
client = AI21Client()
client.chat.completions.create(
messages=messages,
model="jamba-large",
max_tokens=1024,
)
```
```javascript JavaScript theme={"system"}
import { AI21 } from 'ai21';
const client = new AI21({
apiKey: process.env.AI21_API_KEY, // You can also hardcode your API key here
});
const response = await client.chat.completions.create({
model: 'jamba-large',
messages: [
{ role: 'user', content: 'Hello how are you?' }
],
max_tokens: 1024
});
console.log(response);
```
```php PHP theme={"system"}
"https://api.ai21.com/studio/v1/chat/completions",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 30,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "POST",
CURLOPT_POSTFIELDS => "{\n"
. " \"model\": \"jamba-large\",\n"
. " \"messages\": [\n"
. " {\n"
. " \"role\": \"user\",\n"
. " \"content\": \"Hello how are you?\"\n"
. " }\n"
. " ],\n"
. " \"max_tokens\": 1024\n"
. "}",
CURLOPT_HTTPHEADER => [
"Content-Type: application/json",
"Authorization: Bearer YOUR_API_KEY"
],
]);
$response = curl_exec($curl);
$err = curl_error($curl);
curl_close($curl);
if ($err) {
echo "cURL Error #:" . $err;
} else {
echo $response;
```
```bash cURL theme={"system"}
curl https://api.ai21.com/studio/v1/chat/completions \
--header "Authorization: Bearer $AI21_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "jamba-large",
"max_tokens": 1024,
"messages": [
{"role": "user", "content": "Hello, how are you?"}
]
}'
```
```go Go theme={"system"}
package main
import (
"fmt"
"strings"
"net/http"
"io/ioutil"
)
func main() {
url := "https://api.ai21.com/studio/v1/chat/completions"
payload := strings.NewReader("{\n" +
" \"model\": \"jamba-large\",\n" +
" \"messages\": [\n" +
" {\n" +
" \"role\": \"user\",\n" +
" \"content\": \"Hello how are you?\"\n" +
" }\n" +
" ],\n" +
" \"max_tokens\": 1024\n" +
"}")
req, _ := http.NewRequest("POST", url, payload)
req.Header.Add("Content-Type", "application/json")
req.Header.Add("Authorization", "Bearer YOUR_API_KEY")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
body, _ := ioutil.ReadAll(res.Body)
fmt.Println(res)
```
```java Java theme={"system"}
import kong.unirest.HttpResponse;
import kong.unirest.Unirest;
public class AI21Chat {
public static void main(String[] args) {
HttpResponse response = Unirest.post("https://api.ai21.com/studio/v1/chat/completions")
.header("Content-Type", "application/json")
.header("Authorization", "Bearer YOUR_API_KEY")
.body("{\n" +
" \"max_tokens\": 1024,\n" +
" \"messages\": [\n" +
" {\n" +
" \"role\": \"user\",\n" +
" \"content\": \"Hello, how are you?\"\n" +
" }\n" +
" ],\n" +
" \"model\": \"jamba-large\"\n" +
"}")
.asString();
System.out.println(response.getStatus());
System.out.println(response.getBody());
}
}
```
```json JSON theme={"system"}
{
"model": "jamba-large",
"messages": [
{
"role": "user",
"content": "Hello how are you?"
}
],
"max_tokens": 1024
}
```
# Chat response
Source: https://docs.ai21.com/reference/jamba-api-response
## Response details
### Non-streaming results
A successful non-streamed response includes the following members:
Unique ID for each request (not message). Same ID for all responses in a streaming response.
One or more responses, depending on the `n` parameter from the request.
Each response includes the following members.
Zero-based index of the message in the list of messages. Note that this might not correspond with the position in the response list.
The message generated by the model. Includes two fields: `role` and `content`.
Tool calls only occur if a tools parameter was specified in the request. These tool calls apply solely to the current message, and returned values should be added to the message thread in both the assistant message tool\_calls fields and the tool message.
ID of the tool call, generated by the model.
The type of tool called. Currently the only possible value is "function".
The invoked function.
The name of the function, which you specified in your request.
A JSON object containing the function's parameters and values.
Why the message ended.
The response ended naturally as a complete answer (due to end-of-sequence token) or because the model generated a stop sequence provided in the request.
The response ended by reaching max\_tokens.
The token counts for this request.
Per-token billing is based on the prompt token and completion token counts and rates.
Number of tokens in the prompt for this request.
The prompt token contains the entire message history and extra tokens for combining messages, proportional to the number of messages.
Number of tokens in the response message.
prompt\_tokens and completion\_tokens.
### Streamed results
Setting `stream = true` in the request will return a stream of messages, each containing one token. You can read more about streaming calls using the [SDK](https://github.com/AI21Labs/ai21-python/blob/main/README.md#Streaming).
The final message will be `data: [DONE]`. All other messages will have `data` set to a JSON object with the following fields:
An object containing either an object with the following members, or the string "DONE" for the last message.
Unique ID for each request (not message). Same ID for all streaming responses.
An array with one object containing the following fields:
Always zero.
* The first message in the stream will be an object set to `{"role":"assistant"}`.
* Subsequent messages will have an object `{"content": **token**}` with the generated token.
Why the message ended.
The last message includes this field, which shows the total token counts for the request. Per-token billing is based on the prompt token and completion token counts and rates.
When present, it contains a null value except for the last chunk which contains the token usage statistics for the entire request.
Number of tokens in the prompt for this request.
The prompt token contains the entire message history and extra tokens for combining messages, proportional to the number of messages.
Number of tokens in the response message.
prompt\_tokens and completion\_tokens.
`usage` will be `null` except for the last chunk which contains the token usage statistics for the entire request.
```python Python (Non-streaming results) theme={"system"}
import asyncio
from ai21 import AsyncAI21Client
from ai21.models.chat import ChatMessage
messages = [ChatMessage(content="What is the meaning of life?", role="user")]
client = AsyncAI21Client()
async def main():
response = await client.chat.completions.create(
messages=messages,
model="jamba-large",
stream=True,
)
async for chunk in response:
print(chunk.choices[0].delta.content, end="")
asyncio.run(main())
```
```python Python (Streamed results) theme={"system"}
from ai21 import AI21Client
from ai21.models.chat import ChatMessage
messages = [ChatMessage(content="What is the meaning of life?", role="user")]
client = AI21Client()
response = client.chat.completions.create(
messages=messages,
model="jamba-large",
stream=True,
)
for chunk in response:
print(chunk.choices[0].delta.content, end="")
```
***
## Error Codes
500 - Internal Server Error\
429 - Too Many Requests (You are sending requests too quickly.)\
503 - Service Unavailable (The engine is currently overloaded, please try again later)\
401 - Unauthorized (Incorrect API key provided/Invalid Authentication)\
403 - Access Denied\
422 - Unprocurable Entity (Request body is malformed)
***
# Create run
Source: https://docs.ai21.com/reference/maestro-create-run
POST https://api.ai21.com/studio/v1/maestro/runs
## Request body
A list of conversational turns, alternating between the user and the assistant, or a single instruction string.
Please note that there are two different options for the input.
Text input `string`\
A plain instruction string for AI21 Maestro. Equivalent to a single message with the user role.\
\
Input messages `object []`\
A list of message objects representing a multi-turn conversation alternating between the user and the assistant.
The role of the entity that is creating the message. Allowed values include:
`user` : A message is sent by an actual user. It represents the user's intent.
`assistant` : A message is generated by the assistant.
The textual content of the message.
A high\_level instruction that defines AI21 Maestro's overall behavior, role, opreating systems and response style throughout the run.\
\
Common use cases include:\
**Degining identity**: e.g "You are an expert financial analyst". \
**Providing context**: e.g "Today's date is November 10, 2025". \
**Guiding interactions**: e.g "if the user's question is vague, as a clatifying question before answering".
**Deciding whether to use** `requirements` or `system_prompt`\
\
To avoid conflicts and ensure predictable behavior, always prefer using `requirements` for output conditions, and make sure they don’t contradict the `system_prompt`.
`system_prompt` defines the agent’s identity, tone, and core reasoning principles. It shapes how Maestro behaves throughout the entire run.
`requirements`, on the other hand, specify the exact conditions the output must satisfy.
Explicit requirements for the output. The requirements can be in terms of content, style, format, genre, point of view, and various guardrails AI21 Maestro will allocate resources to maintain all requirements throughout the run. At the end of the run, the score for each requirement and the overall score will be returned. Users can specify up to 10 requirements.
The requirement's name, which allows users to reference it in their applications later.
The requirement's description in natural language. This will influence our result scoring and budget allocation. The maximum length for each description is 128 words.
Default: false.\
If set to `true`, and the associated requirement is not fully satisfied, the overall requirements\_result score will be 0.
Examples of requirements:
* "Write up to 3 paragraphs"
* "Output only the answer"
* "Don't mention a word about the company's CEO"
* "Use a friendly tone"
**Avoid conflicting requirements**
Conflicts may reduce run quality or prevent fulfillment.
An array of objects. Each object describes a tool that AI21 Maestro may invoke during execution. The behavior and required fields vary depending on the tool’s `type`.
**MCP Tool** `object`\
Use this tool type to allow AI21 Maestro to interact with external services through the Model Context Protocol (MCP). MCP enables dynamic discovery of available tools and resources provided by an MCP server, making integration more flexible and extensible.
The string has to be `mcp`.
Unique name for the MCP server within the session.
Fully qualified URL of the MCP server endpoint.
Authentication headers (e.g., bearer tokens).
List of tool names that Maestro can use. Must align with the tools exposed by the MCP server.\
If not specified, all tools will be loaded.
**HTTP Tool** `object`\
Use this tool type to allow AI21 Maestro to perform HTTP POST requests to an external service.
The string has to be `http`.
Describes the callable function interface of the tool.
A unique identifier for the tool. Should clearly reflect the action it performs.
An explanation of what the tool does and when it should be used.
A JSON Schema describing the input parameters this tool accepts.\
Each parameter should preferably include a description.
Defines how to construct the HTTP request.
The endpoint URL the tool should call.
HTTP headers to include in the request, typically used for authorization or content-type declarations.
Only `POST`requests are currently supported. \
For more information on how the request is made, see the following example.
**JSON Example**
```json theme={"system"}
{
"type": "http",
"function": {
"name": "getWeatherForecast",
"description": "Fetches a 7-day weather forecast for a given location.",
"parameters": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City or region to fetch weather for"
}
},
"required": ["location"]
}
},
"endpoint": {
"url": "https://api.weatherapi.com/v1/forecast",
"headers": {
"Authorization": "Bearer YOUR_API_KEY"
}
}
}
```
**File Search Tool** `object`\
When provided, this object defines filters that AI21 Maestro will apply whenever it performs a file search.
The string has to be `file_search`.
Restrict file search to files with these labels.\
If not provided, search is not restricted by label.
Restrict file search to these file IDs.\
If not provided, search is not restricted by file ID.
**Web Search Tool** `object`\
When provided, this object defines filters that AI21 Maestro will apply whenever it performs a web search.
The string has to be `web_search`.
Restrict web search to the specified URL prefixes. The following formats are valid:
1. `https://example.com` – limits to URLs that start with `www.example.com`
2. `example.com` – limits to URLs that start with `www.example.com`
3. `example.com/page` – limits to URLs that start with `www.example.com/page`
4. `sub.example.com` – limits to URLs that start with the specified subdomain (e.g., `docs.example.com`, `blog.example.com`)
If not provided, search is not restricted to specific domains.
defaults to null.
Specify a single model to be used for the run. Choose from the following:
**First-Party Models** Models hosted and managed by AI21:
* `jamba-large`
* `jamba-mini`
**Third‑Party Models (Managed by AI21)** Models hosted by third-party providers but accessed through AI21’s API (no user API key required):
* `claude-3-5-haiku`
* `claude-4-sonnet`
* `gemini-2.5-flash`
* `gemini-2.5-flash-lite`
* `gpt-4.1`
* `gpt-4.1-mini`
* `llama-33-70b-instruct`
* `mistral-7b`
* `mistral-8x7b`
* `mistral-small`
* `mistral-small-24B`
* `qwen-qwq-32b`
* `qwen3-235b-a22b-instruct-2507`
**Third-Party Models (BYOK)**\
Bring‑Your‑Own‑Key (BYOK) third‑party models: \
Specify the ID of a third-party model configured on the [**3rd-Party Models page**](https://studio.ai21.com/v2?tab=third_party_models).
These models require your own API key. AI21 uses your key to authenticate and access the model on your behalf. Supported providers include OpenAI, Anthropic, and Google.
Controls how many resources AI21 Maestro allocates to fulfill a task, including the number of execution steps and the level of effort applied.
The higher the budget, the more likely latency and execution costs will increase.\
\
**Accepted values**:
* `low` - One execution step is taken with minimal effort to fulfill the requirements.
* `medium` - Multiple execution steps are taken with moderate effort to fulfill the requirements.
* `high` - Multiple execution strategies, each containing multiple execution steps, are used with maximum effort to fulfill the requirements.
Defaults to `low` if not provided.
Specify which extra fields should be included in the output. If not provided, none of these fields will be included. Supported values:
* `data_sources` — Includes the `data_sources` field in the output.
* `requirements_result` — Includes detailed results for each requirement in the `requirements_result` field.
Controls the output language of AI21 Maestro responses.\
Response language can be only one of the following: "arabic", "dutch", "english", "french", "german", "hebrew", "italian", "portuguese", "spanish".\
\
If not provided, the language is detected automatically, AI21 Maestro will respond in the same language as the input prompt unless a specific language is set.
## Returns
[Run object](/reference/run-object)
```python Python theme={"system"}
from ai21 import AI21Client
from ai21.models.chat import ChatMessage
client = AI21Client(api_key="YOUR_API_KEY")
# Create and poll a Maestro run
run = client.beta.maestro.runs.create_and_poll(
input="Summarize today's major movements in global financial markets and explain possible reasons for these changes.",
requirements=[
{"name": "factual_accuracy",
"description": "Base insights on verifiable data from reliable financial sources."},
{"name": "clarity", "description": "Keep explanations concise and free of jargon."},
{"name": "tone", "description": "Maintain a neutral, professional tone suitable for enterprise reporting."}
],
tools=[
{
"type": "web_search",
"urls": ["https://www.reuters.com/markets/", "https://finance.yahoo.com/"]
}
],
include=["requirements_result", "data_sources"],
budget="medium"
)
# Print the generated analysis
print("=== Financial Market Summary ===")
print(run.result)
print()
# Print requirement evaluation
print("=== Requirements Results ===")
for req in run.requirements_result["requirements"]:
print(f"- {req['name']}: {req['score']} ({req['reason']})")
print()
# Print data sources
print("=== Data Sources ===")
if hasattr(run, 'data_sources') and run.data_sources:
if run.data_sources.get('web_search'):
print("Web Search Results:")
for result in run.data_sources['web_search']:
print(f" - {result.get('url')}")
else:
print("No data sources available")
```
```javascript Javascript theme={"system"}
// TypeScript example using AI21 Labs TypeScript SDK
// Install deps first: npm i ai21 ts-node typescript
// If the import path differs in your SDK version, adjust it per the SDK docs.
import { AI21} from "ai21";
const apiKey = process.env.AI21_API_KEY;
if (!apiKey) {
throw new Error("AI21_API_KEY is not set. Please export it as an environment variable.");
}
const client = new AI21({
apiKey,
});
const body = {
input: [
{
role: "user",
content:
"Summarize today's major movements in global financial markets and explain possible reasons for these changes.",
},
],
requirements: [
{
name: "factual_accuracy",
description:
"Base insights on verifiable data from reliable financial sources.",
},
{
name: "clarity",
description: "Keep explanations concise and free of jargon.",
},
{
name: "tone",
description:
"Maintain a neutral, professional tone suitable for enterprise reporting.",
},
],
tools: [
{
type: "web_search",
urls: [
"https://www.reuters.com/markets/",
"https://finance.yahoo.com/",
],
},
],
budget: "low",
include: ["requirements_result", "data_sources"],
};
async function main() {
const result = await client.beta.maestro.runs.createAndPoll(body, {
interval: 1000,
timeout: 100000
});
console.log(JSON.stringify(result, null, 2));
}
main().catch((err) => {
console.error(err);
process.exit(1);
});
```
```php PHP theme={"system"}
'https://api.ai21.com/studio/v1/maestro/runs',
CURLOPT_RETURNTRANSFER => true,
CURLOPT_FOLLOWLOCATION => true,
CURLOPT_CUSTOMREQUEST => 'POST',
CURLOPT_HTTPHEADER => array(
'Authorization: Bearer YOUR_API_KEY_HERE',
'Content-Type: application/json'
),
CURLOPT_POSTFIELDS => '{
"input": [
{
"role": "user",
"content": "Summarize today\'s major movements in global financial markets and explain possible reasons for these changes."
}
],
"requirements": [
{
"name": "factual_accuracy",
"description": "Base insights on verifiable data from reliable financial sources."
},
{
"name": "clarity",
"description": "Keep explanations concise and free of jargon."
},
{
"name": "tone",
"description": "Maintain a neutral, professional tone suitable for enterprise reporting."
}
],
"tools": [
{
"type": "web_search",
"urls": [
"https://www.reuters.com/markets/",
"https://finance.yahoo.com/"
]
}
]
}'
));
$response = curl_exec($curl);
curl_close($curl);
echo $response;
```
```bash cURL theme={"system"}
curl --location --request POST 'https://api.ai21.com/studio/v1/maestro/runs' \
--header 'Authorization: Bearer ' \
--header 'Content-Type: application/json' \
--data '{
"input": [
{
"role": "user",
"content": "Summarize today'\''s major movements in global financial markets and explain possible reasons for these changes."
}
],
"requirements": [
{
"name": "factual_accuracy",
"description": "Base insights on verifiable data from reliable financial sources."
},
{
"name": "clarity",
"description": "Keep explanations concise and free of jargon."
},
{
"name": "tone",
"description": "Maintain a neutral, professional tone suitable for enterprise reporting."
}
],
"tools": [
{
"type": "web_search",
"urls": [
"https://www.reuters.com/markets/",
"https://finance.yahoo.com/"
]
}
],
"budget": "medium",
"include": ["requirements_result", "data_sources"]
}'
```
```java Java theme={"system"}
HttpResponse response = Unirest.post("https://api.ai21.com/studio/v1/maestro/runs")
.header("Content-Type", "application/json")
.body("{\n \"input\": [\n {\n \"role\": \"\",\n \"content\": \"\"\n }\n ],\n \"requirements\": [\n {\n \"name\": \"\",\n \"description\": \"\"\n }\n ],\n \"tool_resources\": {\n \"file_search\": {\n \"labels\": [\n \"\"\n ],\n \"file_ids\": [\n \"\"\n ]\n },\n \"web_search\": {\n \"urls\": [\n \"\"\n ]\n }\n },\n \"models\": [\n \"\"\n ],\n \"budget\": \"\",\n \"include\": [\n \"\"\n ],\n \"response_language\": {}\n}")
.asString();
```
```json JSON theme={"system"}
{
"input": [
{
"role": "user",
"content": "Summarize today's major movements in global financial markets and explain possible reasons for these changes."
}
],
"requirements": [
{
"name": "factual_accuracy",
"description": "Base insights on verifiable data from reliable financial sources."
},
{
"name": "clarity",
"description": "Keep explanations concise and free of jargon."
},
{
"name": "tone",
"description": "Maintain a neutral, professional tone suitable for enterprise reporting."
}
],
"tools": [
{
"type": "web_search",
"urls": [
"https://www.reuters.com/markets/",
"https://finance.yahoo.com/"
]
}
],
"budget": "medium",
"include": ["requirements_result", "data_sources"]
}
```
# File Library
Source: https://docs.ai21.com/reference/manage-library-ref
## Overview
The library can hold an unlimited number of files, but a maximum of 1GB of files.
Files support the following metadata:
The unique identifier of the file, generated by AI21.
The name of the file specified by you.
An arbitrary file-path-like string you can assign to indicate the content of a file.
The type of the file.
The size of the file in bytes.
The labels associated with the file. You can apply arbitrary string labels to your files and limit queries to files with one or more labels. Similar to paths, but labels do not prefix match. Labels are case-sensitive.
The public URL of the file, specified by you. This URL is not validated by AI21 or used in any way. It is strictly a piece of metadata that you can optionally attach to a file.
The identifier of the user who uploaded the file.
The date when the file was uploaded.
The last update date of the file in the library.
Where was the file uploaded from.
The status of the file in the library.
| Value | Description |
| -------------------------------------- | ------------------------------------------------------------------------- |
| `DB_RECORD_CREATED` | The file record has been created in the database. |
| `UPLOADED` | The file has been uploaded and is being processed. |
| `UPLOAD_FAILED` | The file upload has failed. |
| `PROCESSING` | The file is currently being processed. |
| `PROCESSED` | The file has been processed and is ready for use. |
| `FILE_DOWNLOAD_FAILED` | The file download has failed. |
| `PARSING_FAILED` | Parsing the file has failed. |
| `SEGMENTATION_FAILED` | File segmentation has failed. |
| `EMBEDDING_FAILED` | Embedding the file segments has failed. |
| `VECTORS_DB_UPSERT_FAILED` | Updating the vector database has failed. |
| `VECTOR_DATA_UPLOAD_TO_STORAGE_FAILED` | Uploading the file’s embedding vectors to the storage has failed. |
| `SEGMENTS_UPLOAD_TO_DB_FAILED` | Uploading the file’s segments to the database has failed. |
| `PROCESSING_FAILED` | The processing of the file has failed. |
| `DELETION_FAILED` | The deletion of the file has failed. |
| `RE_PROCESS_CANDIDATE` | The file is a candidate for re-processing. |
| `RE_PROCESSING` | The file is currently being re-processed. |
| `DELETING` | The file is in the process of being deleted. |
| `DELETED_FROM_VECTOR_DB` | The file has been deleted from the vector database. |
| `DELETE_REQUEST_DURING_PROCESSING` | A delete request was received during processing. |
| `MAX_RETRIES_REACHED` | The maximum number of retries has been reached, the operation has failed. |
#### Filtering file results
When you create or update a file, you can optionally apply `path` and/or `label` values to a file. You can later search, query, or list files based on matching paths or labels. This can be useful to query subsets of documents.
For example, you can assign a path of either `/financial/USA` or `/financial/UK` to certain documents and later query only US financial documents, only UK financial documents, or query all financial documents by specifying `/financial/` as the path.
# Delete file
Source: https://docs.ai21.com/reference/manage-library-ref/delete-file
DELETE https://api.ai21.com/studio/v1/library/files/{file_id}
Delete the specified file from the library.
**Restrictions**:
Files in `PROCESSING` status cannot be deleted. Attempts to delete such files will result in a 422 error.
## Responses
### Response 200
Successful deletion. No body content.
### Response 422
File by this ID does not exist.
```python Python theme={"system"}
from ai21 import AI21Client
# Initialize the client
client = AI21Client(
api_key="your_api_key_here" # or use environment variable AI21_API_KEY
)
# Delete a file by its ID
try:
client.library.files.delete("your-file-id-here")
print("File deleted successfully")
except Exception as e:
print(f"Error deleting file: {e}")
```
```javascript JavaScript theme={"system"}
const axios = require('axios');
const apiKey = 'your_api_key_here';
const fileId = 'your-file-id-here';
axios.delete(`https://api.ai21.com/studio/v1/library/files/${fileId}`, {
headers: {
'Authorization': `Bearer ${apiKey}`
}
})
.then(() => console.log('File deleted successfully'))
.catch(error => console.error('Error deleting file:', error.message));
```
```php PHP theme={"system"}
```
```go Go theme={"system"}
package main
import (
"fmt"
"net/http"
)
func main() {
apiKey := "your_api_key_here"
fileId := "your-file-id-here"
req, _ := http.NewRequest("DELETE", "https://api.ai21.com/studio/v1/library/files/"+fileId, nil)
req.Header.Set("Authorization", "Bearer "+apiKey)
client := &http.Client{}
resp, err := client.Do(req)
if err != nil {
fmt.Println("Error deleting file:", err)
} else {
fmt.Println("File deleted successfully")
resp.Body.Close()
}
}
```
```bash cURL theme={"system"}
curl -X DELETE https://api.ai21.com/studio/v1/library/files/your-file-id-here \
-H "Authorization: Bearer your_api_key_here"
```
```java Java theme={"system"}
import okhttp3.*;
public class DeleteFile {
public static void main(String[] args) throws Exception {
String apiKey = "your_api_key_here";
String fileId = "your-file-id-here";
OkHttpClient client = new OkHttpClient();
Request request = new Request.Builder()
.url("https://api.ai21.com/studio/v1/library/files/" + fileId)
.delete()
.addHeader("Authorization", "Bearer " + apiKey)
.build();
try (Response response = client.newCall(request).execute()) {
if (response.isSuccessful()) {
System.out.println("File deleted successfully");
} else {
System.out.println("Error deleting file: " + response.message());
}
}
}
}
```
# Get file download link
Source: https://docs.ai21.com/reference/manage-library-ref/get-file-download-link
GET https://api.ai21.com/studio/v1/library/files/{file_id}/download
Retrieve a signed URL you can use to download a file’s contents from your workspace library.
## Path parameters
The unique ID of the file to download.
## Response 200
Returns a string containing the signed download URL for the requested file.
Example response:
```json theme={"system"}
"https://storage.ai21.com/files/file_123abc/download?token=xyz"
```
### Response 422
No matching ID is found.
```python Python theme={"system"}
import requests
file_id = "your_file_id_here"
url = f"https://api.ai21.com/studio/v1/library/files/{file_id}/download"
headers = {"Authorization": "Bearer your_api_key_here"}
response = requests.get(url)
if response.status_code == 200:
print("Download URL:", response.text)
else:
print("Error:", response.status_code, response.text)
```
```javascript JavaScript theme={"system"}
const file_id = "your_file_id_here";
fetch(`https://api.ai21.com/studio/v1/library/files/${file_id}/download`, {
method: "GET",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
},
})
.then(res => res.text())
.then(link => console.log("Download URL:", link))
.catch(err => console.error("Error:", err));
```
```php PHP theme={"system"}
"https://api.ai21.com/studio/v1/library/files/$file_id/download",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 30,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "GET",
CURLOPT_HTTPHEADER => [
"Authorization: Bearer YOUR_API_KEY"
],
]);
$response = curl_exec($curl);
$err = curl_error($curl);
curl_close($curl);
if ($err) {
echo "cURL Error #:" . $err;
} else {
echo "Download URL: " . $response;
}
```
```go Go theme={"system"}
package main
import (
"fmt"
"io/ioutil"
"net/http"
)
func main() {
fileID := "your_file_id_here"
url := fmt.Sprintf("https://api.ai21.com/studio/v1/library/files/%s/download", fileID)
req, err := http.NewRequest("GET", url, nil)
if err != nil {
fmt.Println("Error creating request:", err)
return
}
req.Header.Add("Authorization", "Bearer YOUR_API_KEY")
client := &http.Client{}
resp, err := client.Do(req)
if err != nil {
fmt.Println("Error sending request:", err)
```
```cURL theme={"system"}
curl -X GET "https://api.ai21.com/studio/v1/library/files/{file_id}/download" \
-H "Authorization: Bearer YOUR_API_KEY"
```
# Get workspace file
Source: https://docs.ai21.com/reference/manage-library-ref/get-workspace-files
GET https://api.ai21.com/studio/v1/library/files/{file_id}
Retrieve metadata about a specific file.
The `file_id` is generated by AI21 when you upload the file.
## Path Parameters
The unique ID of the file to retrieve.
## Responses
### Response 200
A successful response returns metadata about the requested file.
The file ID.
The file name.
File size in bytes.
File creation timestamp.
Labels assigned to the file
Error code, if any.
Error message, if any.
### Response 422
No matching ID is found.
```python Python theme={"system"}
from ai21 import AI21Client
client = AI21Client()
# Replace with your actual file_id
file_id = "file_123abc"
# Retrieve file metadata
file_metadata = client.library.files.get(file_id=file_id)
print(file_metadata)
```
```javascript JavaScript theme={"system"}
const options = {method: 'GET'};
fetch('https://api.ai21.com/studio/v1/library/files/{file_id}', options)
.then(response => response.json())
.then(response => console.log(response))
.catch(err => console.error(err));
```
```php PHP theme={"system"}
"https://api.ai21.com/studio/v1/library/files/{file_id}",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 30,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "GET",
]);
$response = curl_exec($curl);
$err = curl_error($curl);
curl_close($curl);
if ($err) {
echo "cURL Error #:" . $err;
} else {
echo $response;
}
```
```go Go theme={"system"}
package main
import (
"fmt"
"net/http"
"io/ioutil"
)
func main() {
url := "https://api.ai21.com/studio/v1/library/files/{file_id}"
req, _ := http.NewRequest("GET", url, nil)
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
body, _ := ioutil.ReadAll(res.Body)
fmt.Println(res)
fmt.Println(string(body))
}
```
```bash cURL theme={"system"}
curl --request GET \
--url https://api.ai21.com/studio/v1/library/files/{file_id}
```
```java Java theme={"system"}
HttpResponse response = Unirest.get("https://api.ai21.com/studio/v1/library/files/{file_id}")
.asString();
```
```json JSON theme={"system"}
{
"id": "file_123abc",
"name": "customer_data.pdf",
"size": 204800,
"created_at": "2025-10-20T14:23:11Z",
"labels": ["invoices", "Q3"]
}
```
# List library files
Source: https://docs.ai21.com/reference/manage-library-ref/list-library-files
GET https://api.ai21.com/studio/v1/library/files
Retrieve a list of documents in the user's library. You can optionally specify a filter to return only files with matching labels.
When specifying qualifiers with your request, only files that match all qualifiers will be returned. For example, if you specify `label='financial'` and `status='UPLOADED'`, only files with the label 'financial' **and** status 'UPLOADED' will be returned.
## Request parameters
The following optional parameters can be used to filter your results.
The full name of the uploaded file, without any path parameters. So: "mydoc.txt", not "/users/benfranklyn/Documents/mydoc.txt". Does not match name substrings, so "doc.txt" does not match "mydoc.txt".
Status of the file in the library. Supported values:
* `DB_RECORD_CREATED`
* `UPLOADED`
* `UPLOAD_FAILED`
* `PROCESSED`
* `PROCESSING_FAILED`
Return only files with this label. Label matching is case-sensitive, and will not match substrings.
By default, the endpoint returns up to 10000 files. Pagination can be controlled using the following parameters:
The number of files to skip.
The number of files to retrieve (maximum 1,000, default 1,000)
## Responses
### Response 200
A successful response returns an array of file metadata items.
```python Python theme={"system"}
import requests
url = "https://api.ai21.com/studio/v1/library/files"
payload = {
"name": "",
"status": "",
"label": [""],
"Pagination": {
"offset": 123,
"limit": 123
}
}
headers = {"Content-Type": "application/json"}
response = requests.request("GET", url, json=payload, headers=headers)
print(response.text)
```
```javascript JavaScript theme={"system"}
const options = {
method: 'GET',
headers: {'Content-Type': 'application/json'},
body: '{"name":"","path":"","status":"","label":[""],"Pagination":{"offset":123,"limit":123}}'
};
fetch('https://api.ai21.com/studio/v1/library/files', options)
.then(response => response.json())
.then(response => console.log(response))
.catch(err => console.error(err));
```
```php PHP theme={"system"}
"https://api.ai21.com/studio/v1/library/files",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 30,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "GET",
CURLOPT_POSTFIELDS => "{\n \"name\": \"\",\n \"path\": \"\",\n \"status\": \"\",\n \"label\": [\n \"\"\n ],\n \"Pagination\": {\n \"offset\": 123,\n \"limit\": 123\n }\n}",
CURLOPT_HTTPHEADER => [
"Content-Type: application/json"
],
]);
$response = curl_exec($curl);
$err = curl_error($curl);
curl_close($curl);
```
```go Go theme={"system"}
package main
import (
"fmt"
"strings"
"net/http"
"io/ioutil"
)
func main() {
url := "https://api.ai21.com/studio/v1/library/files"
payload := strings.NewReader("{\n \"name\": \"\",\n \"path\": \"\",\n \"status\": \"\",\n \"label\": [\n \"\"\n ],\n \"Pagination\": {\n \"offset\": 123,\n \"limit\": 123\n }\n}")
req, _ := http.NewRequest("GET", url, payload)
req.Header.Add("Content-Type", "application/json")
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
body, _ := ioutil.ReadAll(res.Body)
fmt.Println(res)
fmt.Println(string(body))
}
```
```bash cURL theme={"system"}
curl --request GET \
--url https://api.ai21.com/studio/v1/library/files \
--header 'Content-Type: application/json' \
--data '{
"name": "",
"path": "",
"status": "",
"label": [
""
],
"Pagination": {
"offset": 123,
"limit": 123
}
```
```java Java theme={"system"}
HttpResponse response = Unirest.get("https://api.ai21.com/studio/v1/library/files")
.header("Content-Type", "application/json")
.queryString("name", "")
.queryString("path", "")
.queryString("status", "")
.queryString("label", "")
.queryString("offset", 123)
.queryString("limit", 123)
.asString();
```
# Update file
Source: https://docs.ai21.com/reference/manage-library-ref/update-file
PUT https://api.ai21.com/studio/v1/library/files/{file_id}
Update the specified parameters of a specific document in the user's library. This operation currently supports updating the publicUrl and labels parameters.
## Request body
The updated public URL of the document.
The updated labels associated with the file. Separate multiple labels with commas.
## Responses
If the update is successful, the server will return an HTTP status code 200 with no body content. If the document ID does not exist, a 422 error message will be returned.
```python Python theme={"system"}
from ai21 import AI21Client
client = AI21Client()
file_id = client.library.files.create(
labels=["label1", "label2"],
public_url="www.example.com",
)
uploaded_file = client.library.files.get(file_id)
```
```javascript JavaScript theme={"system"}
const options = {
method: 'PUT',
headers: {'Content-Type': 'application/json'},
body: '{"publicUrl":"","labels":[""]}'
};
fetch('https://api.ai21.com/studio/v1/library/files/{file_id}', options)
.then(response => response.json())
.then(response => console.log(response))
.catch(err => console.error(err));
```
```php PHP theme={"system"}
"https://api.ai21.com/studio/v1/library/files/{file_id}",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 30,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "PUT",
CURLOPT_POSTFIELDS => "{\n \"publicUrl\": \"\",\n \"labels\": [\n \"\"\n ]\n}",
CURLOPT_HTTPHEADER => [
"Content-Type: application/json"
],
]);
$response = curl_exec($curl);
$err = curl_error($curl);
curl_close($curl);
if ($err) {
echo "cURL Error #:" . $err;
} else {
echo $response;
}
```
```go Go theme={"system"}
package main
import (
"fmt"
"strings"
"net/http"
"io/ioutil"
)
func main() {
url := "https://api.ai21.com/studio/v1/library/files/{file_id}"
payload := strings.NewReader("{\n \"publicUrl\": \"\",\n \"labels\": [\n \"\"\n ]\n}")
req, _ := http.NewRequest("PUT", url, payload)
req.Header.Add("Content-Type", "application/json")
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
body, _ := ioutil.ReadAll(res.Body)
fmt.Println(res)
fmt.Println(string(body))
}
```
```bash cURL theme={"system"}
curl --request PUT \
--url https://api.ai21.com/studio/v1/library/files/{file_id} \
--header 'Content-Type: application/json' \
--data '{
"publicUrl": "",
"labels": [
""
]
}'
```
```java Java theme={"system"}
HttpResponse response = Unirest.put("https://api.ai21.com/studio/v1/library/files/{file_id}")
.header("Content-Type", "application/json")
.body("{\n \"publicUrl\": \"\",\n \"labels\": [\n \"\"\n ]\n}")
.asString();
```
# Upload workspace files
Source: https://docs.ai21.com/reference/manage-library-ref/upload-workspace-files
POST https://api.ai21.com/studio/v1/library/files
Upload files to use.
You can assign metadata to your files to limit searches to specific files by file metadata. There is no bulk upload method; files must be loaded one at a time.
* **Max number of files:** No limit. The playground limits bulk uploads to 50 files per request.
* **Max total library size:** 1 GB.
* **Max file size:** 5 MB.
* **Supported file types:** PDF, DocX, HTML, TXT, Markdown.
## Request body fields
Raw file bytes. Every uploaded file must have a unique file name. Specifying file labels will find only files with *both* the specified name and label.
Arbitrary string labels that describe the contents of this file. Labels are case-sensitive. Specifying labels or labels + name will return only files with both any of the specified labels and the specified name.
A public URL associated with the file, if any. Only used as metadata, to indicate the location of the source file. For example, if implementing a search engine against a website, specifying a URL for each uploaded file is a simple way to present the link to the file in the search results presented to the user.
## Responses
### Response 200
A unique identifier for the uploaded file. Use this later to request, modify, or delete the file. You don't need to store the value though, as it is returned along with all file information in a GET /files request. **Type:** UUID. Example: da13301a-14e4-4487-aa2f-cc6048e73cdc // file-uuid
### Error: Unsupported document type 422
Error message:
```json theme={"system"}
{
"detail": "Invalid file type: image/png. Supported file types are: text/plain, text/html, application/docx, application/pdf"
}
```
```json 200 theme={"system"}
{
"id": "da13301a-14e4-4487-aa2f-cc6048e73cdc",
}
```
```json 422 Unsupported document type theme={"system"}
{
"detail": "Invalid file type: image/png. Supported file types are: text/plain, text/html, application/docx, application/pdf"
}
```
```python Python theme={"system"}
rom ai21 import AI21Client
client = AI21Client()
# Option 1: Upload a local file
file_from_disk = client.library.files.create(
labels=["label1", "label2"],
)
# Option 2: Upload a file from a public URL
file_from_url = client.library.files.create(
public_url="https://www.example.com/file.pdf",
labels=["label1", "label2"],
)
# Retrieve uploaded file metadata (example)
uploaded_file = client.library.files.get(file_from_disk.id)
```
```javascript JavaScript theme={"system"}
const payload = {
path: "docs/file-from-url.pdf",
labels: ["label1", "label2"],
publicUrl: "https://www.example.com/file.pdf"
};
fetch("https://api.ai21.com/studio/v1/library/files", {
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json"
},
body: JSON.stringify(payload)
})
.then(response => response.json())
.then(response => console.log(response))
.catch(error => console.error(error));
```
```php PHP theme={"system"}
"docs/file-from-url.pdf",
"labels" => ["label1", "label2"],
"publicUrl" => "https://www.example.com/file.pdf"
]);
$curl = curl_init();
curl_setopt_array($curl, [
CURLOPT_URL => "https://api.ai21.com/studio/v1/library/files",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_ENCODING => "",
CURLOPT_MAXREDIRS => 10,
CURLOPT_TIMEOUT => 30,
CURLOPT_HTTP_VERSION => CURL_HTTP_VERSION_1_1,
CURLOPT_CUSTOMREQUEST => "POST",
CURLOPT_POSTFIELDS => $payload,
CURLOPT_HTTPHEADER => [
"Authorization: Bearer YOUR_API_KEY",
"Content-Type: application/json"
],
]);
$response = curl_exec($curl);
$err = curl_error($curl);
curl_close($curl);
if ($err) {
echo "cURL Error #:" . $err;
} else {
echo $response;
}
```
```go Go theme={"system"}
package main
import (
"bytes"
"encoding/json"
"fmt"
"io/ioutil"
"net/http"
)
func main() {
url := "https://api.ai21.com/studio/v1/library/files"
payload := map[string]interface{}{
"path": "docs/file-from-url.pdf",
"labels": []string{"label1", "label2"},
"publicUrl": "https://www.example.com/file.pdf",
}
jsonData, err := json.Marshal(payload)
if err != nil {
panic(err)
}
req, err := http.NewRequest("POST", url, bytes.NewBuffer(jsonData))
if err != nil {
panic(err)
}
req.Header.Add("Content-Type", "application/json")
req.Header.Add("Authorization", "Bearer YOUR_API_KEY")
client := &http.Client{}
res, err := client.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
body, err := ioutil.ReadAll(res.Body)
if err != nil {
panic(err)
```
```bash cURI theme={"system"}
curl --request POST \
--url https://api.ai21.com/studio/v1/library/files \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer YOUR_API_KEY' \
--data '{
"path": "docs/file-from-url.pdf",
"labels": ["label1", "label2"],
"publicUrl": "https://www.example.com/file.pdf"
}'
```
```java Java theme={"system"}
HttpResponse response = Unirest.post("https://api.ai21.com/studio/v1/library/files")
.header("Content-Type", "application/json")
.header("Authorization", "Bearer YOUR_API_KEY")
.body("{"
+ "\"path\": \"docs/file-from-url.pdf\","
+ "\"labels\": [\"label1\", \"label2\"],"
+ "\"publicUrl\": \"https://www.example.com/file.pdf\""
+ "}")
.asString();
```
# Retrieve run status
Source: https://docs.ai21.com/reference/retrieve-run
GET https://api.ai21.com/studio/v1/maestro/runs/{run_id}
Retrieves the status of the run execution based on its ID. The status can be `completed`, `failed`, or `in_progress`.
## Path parameters
A unique identifier of the run you wish to retrieve.
## Returns
The [run object](/reference/run-object) matches the specified ID.
```python Python theme={"system"}
from ai21 import AI21Client
client = AI21Client(api_key="your-api-key") # Replace with your API key
run_id = "your_run_id_here"
# Retrieve the Maestro run by ID
retrieved_run = client.beta.maestro.runs.retrieve(run_id)
print(retrieved_run)
```
```javascript JavaScript theme={"system"}
const runId = "your_run_id_here";
fetch(`https://api.ai21.com/studio/v1/maestro/runs/${runId}`, {
method: "GET",
headers: {
"Authorization": "Bearer your-api-key",
"Content-Type": "application/json"
}
})
.then(res => res.json())
.then(data => console.log(data))
.catch(err => console.error(err));
```
```php PHP theme={"system"}
"https://api.ai21.com/studio/v1/maestro/runs/$run_id",
CURLOPT_RETURNTRANSFER => true,
CURLOPT_CUSTOMREQUEST => "GET",
CURLOPT_HTTPHEADER => [
"Authorization: Bearer your-api-key",
"Content-Type: application/json"
],
]);
$response = curl_exec($curl);
$error = curl_error($curl);
curl_close($curl);
if ($error) {
echo "cURL Error: " . $error;
} else {
echo $response;
}
```
```go Go theme={"system"}
package main
import (
"fmt"
"io/ioutil"
"net/http"
)
func main() {
runID := "your_run_id_here"
url := fmt.Sprintf("https://api.ai21.com/studio/v1/maestro/runs/%s", runID)
req, err := http.NewRequest("GET", url, nil)
if err != nil {
panic(err)
}
req.Header.Add("Authorization", "Bearer your-api-key")
req.Header.Add("Content-Type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
body, err := ioutil.ReadAll(res.Body)
if err != nil {
panic(err)
}
fmt.Println("Status:", res.Status)
fmt.Println("Response:", string(body))
}
```
```bash cURL theme={"system"}
curl --request GET \
--url https://api.ai21.com/studio/v1/maestro/runs/{run_id} \
--header "Authorization: Bearer your-api-key" \
--header "Content-Type: application/json"
```
```java Java theme={"system"}
String runId = "your_run_id_here";
HttpResponse response = Unirest.get("https://api.ai21.com/studio/v1/maestro/runs/" + runId)
.header("Authorization", "Bearer your-api-key")
.header("Content-Type", "application/json")
.asString();
System.out.println(response.getBody());
```
# Run object
Source: https://docs.ai21.com/reference/run-object
## Run Object
ID of the task that a user can use for polling its status
any of `completed`, `failed`, `in_progress`
The final result object, may vary based on task type.
Contains the retrieved contextual information used during the run.\
Includes optional `web_search`, `file_search`, and `tool_calls` result arrays.
A list of web search results retrieved during the run.
The URL address returned by the search.
The extracted text from the URL that was used as context.
Relevancy score from 0 to 1.
A list of file search results retrieved during the run.
The ID of the file that contains relevant text.
The name of the retrieved file.
The retrieved text segment.\
When `null`, it indicates that the entire file was retrieved.
Relevancy score from 0 to 1.
The order of the retrieved segments when multiple results come from the same file.
A list of tool call results executed during the run.
The name of the tool that was called.
The type of tool invoked: `"mcp"` or `"http"`.
When the tool belongs to an MCP server, indicates which server provided the tool.
A JSON object representing the parameters sent to the tool.
The JSON response returned by the tool.
Detailed results for each requirement.
A value between 0 and 1 indicating how well the requirement was fulfilled.
A textual reason what happened - one of `Perfect score` or `Budget exhausted`.
A list of search requirement shown in the report.
The name provided (or default name) of the requirement requested to be fulfilled
When a user provides the requirement explicitly, it's what the user provided
A value between 0 and 1 indicating how well the requirement was fulfilled.
A textual reason about why the requirement didn't reach a perfect score.
An object that provides details about why the run failed. Returns `null` if the run completed successfully.
A descriptive message explaining the reason for the failure.
The example below shows a fully completed run object, including the `requirements_result`, `data_sources`, and `error` fields.
These fields are **not included by default** and will only appear in the response if you explicitly request them using the include parameter.
For details on how to use this parameter, see the `include` [parameter documentation](https://docs.ai21.com/reference/endpoints#param-include).
### Example:
```json theme={"system"}
{
"id": "dba286ef-5067-4c6e-b215-5483500da8df",
"status": "completed",
"result": "- example 1\n- example 2\n- example 3",
"requirements_result": {
"score": 1,
"finish_reason": "Perfect result achieved",
"requirements": [
{
"name": "bullets",
"description": "Return a bulleted list where each field is a bullet",
"is_mandatory": false,
"score": 1,
"reason": "None"
}
]
},
"data_sources": {
"web_search": [
{
"id": "0686b7a3-6cb6-7128-8000-146778a91763",
"text": "web search result text",
"url": "https://example.com",
"score": 0.01461567
}
],
"file_search": [
{
"text": "file search result text",
"file_id": "aa50de3e-1624-4e9f-a6f4-488a1aa16759",
"file_name": "example.pdf",
"score": 0.7331293
}
]
},
"error": {
"message": "The run failed due to an unexpected server error."
}
}
```