Skip to main content
Version: v2.0

Models

Models are the foundation of intelligence in the aixplain ecosystem. They encapsulate capabilities such as language understanding, translation, summarisation, speech processing, and more. Models can be run directly or integrated as tools inside agents.

All models share a consistent SDK interface — swap one for another without rewriting pipelines.

Setup​

pip install aixplain
from aixplain import Aixplain

aix = Aixplain(api_key="YOUR_API_KEY")

Quick start​

model = aix.Model.get("openai/gpt-4o")

response = model.run(text="Why did the chicken cross the road?")
print(response.data)
Show output

Discovering models​

# By keyword
llama_models = aix.Model.search("llama")["results"]

# By host, developer, or vendor
openai_models = aix.Model.search("", host="openai")["results"]
meta_models = aix.Model.search("", developer="meta")["results"]
anthropic_models = aix.Model.search("", vendor="anthropic")["results"]

for model in openai_models[:5]:
print(model.name, model.id)
Show output

search() returns a Page, not a dict. Reach the items with .results, with ["results"], or by iterating the page directly — all three work, so the ["results"] form used above and the .results form used in the Marketplace reference are equivalent. The page also carries total, page_number, and page_total.

Search parameters:

ParameterDescription
queryKeyword matched against name and description
hostHosting platform (e.g. "openai", "groq")
developerModel developer (e.g. "meta")
vendorModel vendor / supplier (e.g. "anthropic")

Get a specific model​

# By path
model = aix.Model.get("openai/gpt-4o")

# By ID
model = aix.Model.get("6646261c6eb563165658bbb1")

print(model.name, model.id, model.host)
Show output

Running models​

Synchronous​

model = aix.Model.get("openai/gpt-4o")

response = model.run(text="Explain quantum computing in simple terms")
print(response.data)
print(response.status)
Show output

Supported input types​

TypeExample
Textmodel.run(text="Your prompt") or model.run(data="file.txt")
Imagemodel.run(data="image.png")
Audiomodel.run(data="audio.wav")
Videomodel.run(data="video.mp4")
Structuredmodel.run({"text": "prompt", "context": "..."})
note

Image, audio, and video are not supported on-prem. Format and size limits vary by vendor — check the model's page in Studio.

Streaming​

Use stream=True (or run_stream()) to receive response chunks as they are generated instead of waiting for the full output.

model = aix.Model.get("openai/gpt-4o")

with model.run_stream(text="Tell me a short story about a robot.") as stream:
for chunk in stream:
print(chunk.data, end="", flush=True)
Show output

Each chunk has the following fields:

FieldTypeDescription
datastrText content of this chunk
statusstrChunk status ("SUCCESS", "IN_PROGRESS", etc.)
finish_reasonstr | NoneWhy generation stopped ("stop", "length", etc.)
usageobject | NoneToken usage — available on the final chunk

You can also pass stream=True directly to run():

for chunk in model.run(text="Summarise the water cycle.", stream=True):
print(chunk.data, end="", flush=True)
Show output
note

Not all models support streaming. Check model.supports_streaming before calling run_stream() — a ValidationError is raised if the model does not support it.

Asynchronous​

Use run_async() when you don't want to block on long-running tasks:

import time

model = aix.Model.get("openai/gpt-4o")

start = model.run_async(text="Summarise the history of computing.")

while True:
if not start.url: # task finished immediately
print(start.data)
break

result = model.poll(start.url)

if result.completed:
print(result.data)
break

time.sleep(5)
Show output

Batch async​

Start multiple tasks in parallel, then poll until all complete:

import time

model = aix.Model.get("openai/gpt-4o")

prompts = [
"Summarise the benefits of cloud computing",
"Explain blockchain technology",
"Describe machine learning applications",
]

# Kick off all tasks
pending_urls = []
for prompt in prompts:
start = model.run_async(text=prompt)
if start.url:
pending_urls.append(start.url)
else:
print("Completed immediately:", start.data)

# Poll until all finish
results = []
while pending_urls:
for url in pending_urls[:]:
result = model.poll(url)
if result.completed:
results.append(result.data)
pending_urls.remove(url)
time.sleep(3)

for i, output in enumerate(results):
print(f"\n=== Result {i+1} ===\n{output}")
Show output

Configuring model parameters​

Use model.inputs to set, inspect, and reset parameters before calling run(). The proxy supports dot notation, dict notation, and bulk updates.

model = aix.Model.get("openai/gpt-4o")

# Set parameters
model.inputs.temperature = 0.3 # dot notation
model.inputs['max_tokens'] = 1024 # dict notation
model.inputs.update(temperature=0.2, max_tokens=1200) # bulk

# Inspect
print(model.inputs.keys()) # all parameter names
print(model.inputs.required) # required-only
print(dict(model.inputs.items())) # current values as dict

# Reset
model.inputs.reset("temperature") # single param
model.inputs.reset() # all params
Show output
note

Models do not have .actions. Use model.inputs directly to configure parameters.

# Run with configured parameters
model.inputs.temperature = 0.3
model.inputs.max_tokens = 1024

response = model.run(text="Generate a product description")
print(response.data)
print(f"Completion tokens: {response.usage.completion_tokens}")
Show output

Common LLM parameters:

ParameterDescriptionRange
temperatureRandomness — lower is more deterministic0.0 – 2.0
max_tokensMaximum output length1 – model limit
top_pNucleus sampling threshold0.0 – 1.0
frequency_penaltyReduces token repetition-2.0 – 2.0
presence_penaltyEncourages new topics-2.0 – 2.0

Temperature guidance:

model.inputs.temperature = 0.9   # creative tasks (stories, brainstorming)
model.inputs.temperature = 0.3 # factual tasks (summaries, analysis)
model.inputs.temperature = 0.0 # deterministic tasks (code, math)

Get the provider's raw output​

includeRawData is a model-type-agnostic option — it works the same for LLMs, speech, vision, and any other model. Pass options={"includeRawData": True} to receive the backing provider's full, unmodified response alongside the normalized result. The shape of that raw payload mirrors whatever the supplier returns, so it varies by model and provider (for Whisper, for example, it includes per-segment timestamps, token IDs, and log-probs — see Speech recognition).

result = model.run(text="Summarize this article...", options={"includeRawData": True})

print(result.data) # normalized output
print(result._raw_data["rawData"]) # provider's full raw response

Over REST, add "options": {"includeRawData": true} to the request body; the raw payload comes back in a top-level rawData field.

Using models with agents​

Pass a model as the llm for an agent's reasoning loop, or attach it as a tool:

llm = aix.Model.get("openai/gpt-4o")

agent = aix.Agent(
name="Research Assistant",
description="Answers research questions thoroughly.",
instructions="Use the provided LLM to answer questions accurately.",
llm=llm,
)
agent.save()

response = agent.run(query="Explain the difference between supervised and unsupervised learning.")
print(response.data.output)
Show output

Advanced examples​

Compare two models​

gpt4   = aix.Model.get("openai/gpt-4o")
claude = aix.Model.get("anthropic/claude-3-5-sonnet-v2")

prompt = "Explain the concept of recursion in programming."

print("=== GPT-4o ===")
print(gpt4.run(text=prompt).data)

print("\n=== Claude ===")
print(claude.run(text=prompt).data)
Show output

Translation​

translator = aix.Model.get("google/cloud-translation")

response = translator.run(
text="Hello, how are you?",
sourcelanguage="en",
targetlanguage="es",
)
print(response.data)
Show output

Speech recognition​

Whisper Large (66311fda6eb563279c574b71) transcribes audio to text and auto-detects the spoken language — pass any audio URL and get back a full transcript with per-segment timestamps.

Python SDK:

from aixplain import Aixplain

aix = Aixplain(api_key="YOUR_API_KEY")

model = aix.Model.get("66311fda6eb563279c574b71")
result = model.run(
source_audio="https://dare.wisc.edu/wp-content/uploads/sites/1051/2008/04/Arthur.mp3",
sourcelanguage="en",
options={"includeRawData": True},
)
print(result.data)
Show output

curl:

curl --location 'https://models.aixplain.com/api/v2/execute/66311fda6eb563279c574b71' \
--header 'Content-Type: application/json' \
--header 'x-api-key: YOUR_API_KEY' \
--data '{
"source_audio": "https://dare.wisc.edu/wp-content/uploads/sites/1051/2008/04/Arthur.mp3",
"options": {"includeRawData": true}
}'
Show output

Step 1: Get the model​

model = aix.Model.get("66311fda6eb563279c574b71")
print(model.name, model.id)
Show output

Step 2: Upload a local file​

Use FileUploader to push a local audio file to temporary storage, then pass the returned URL to model.run():

from aixplain.v2 import FileUploader

uploader = FileUploader(api_key="YOUR_API_KEY")
audio_url = uploader.upload(
file_path="/path/to/audio.mp3",
is_temp=True,
return_download_link=True,
)

result = model.run(
source_audio=audio_url,
sourcelanguage="en", # required field — model auto-detects actual language
options={"includeRawData": True},
)

print("Transcript:", result.data)
print("Detected language:", result._raw_data["rawData"]["language"])
print("Duration (s):", result._raw_data["rawData"]["duration"])
Show output

sourcelanguage is a required field but its value does not affect language detection — Whisper identifies the spoken language automatically.

Not using the SDK?

FileUploader wraps a presigned-S3 upload you can call with plain REST/curl, then pass the returned URL to any model or agent. See Upload a file via REST for the three-step flow and per-type upload size limits (audio 50 MB, image/documents 25 MB, video/database 300 MB).

Step 3: Read timestamped segments​

Pass options={"includeRawData": True} to get Whisper's full provider response, including per-segment timestamps and confidence scores:

segments = result._raw_data["rawData"]["segments"]
print(f"Total segments: {len(segments)}")

for seg in segments[:3]:
print(f"[{seg['start']:.2f}s → {seg['end']:.2f}s] {seg['text'].strip()}")
Show output

Each segment includes avg_logprob for confidence and no_speech_prob for silence detection.

Step 4: Call via REST API​

Skip the SDK entirely when you already have a public audio URL:

import requests

MODEL_ID = "66311fda6eb563279c574b71"

response = requests.post(
f"https://models.aixplain.com/api/v2/execute/{MODEL_ID}",
headers={"Content-Type": "application/json", "x-api-key": "YOUR_API_KEY"},
json={
"source_audio": "https://dare.wisc.edu/wp-content/uploads/sites/1051/2008/04/Arthur.mp3",
"options": {"includeRawData": True},
},
)
data = response.json()

print("Status:", data.get("status"))
print("Language:", data.get("rawData", {}).get("language"))
print("Transcript:", data.get("data", "")[:120])
Show output

Equivalent curl:

curl --location 'https://models.aixplain.com/api/v2/execute/66311fda6eb563279c574b71' \
--header 'Content-Type: application/json' \
--header 'x-api-key: YOUR_API_KEY' \
--data '{
"source_audio": "https://dare.wisc.edu/wp-content/uploads/sites/1051/2008/04/Arthur.mp3",
"options": {"includeRawData": true}
}'

Parameters:

ParameterTypeRequiredDescription
source_audiostr✅URL to the audio file. Must be publicly accessible.
sourcelanguagestr✅Required field — use "en". Does not restrict language detection.
options.includeRawDatabool—Returns Whisper's full provider response: segments, log-probs, and duration.

Temperature sweep​

model = aix.Model.get("openai/gpt-4o")
prompt = "Complete this sentence: The future of AI is"

for temp in [0.0, 0.5, 1.0, 1.5]:
model.inputs.temperature = temp
print(f"\nTemperature {temp}:", model.run(text=prompt).data)
Show output

Troubleshooting​

Model not found - Verify the path or ID with aix.Model.search(). Check that your API key has access to the model.

Invalid parameters - Confirm supported parameters and valid ranges on the model's Studio page. Not all models accept all parameters.

Async tasks timing out - Increase the time.sleep() interval for long-running tasks. Check the aixplain dashboard to confirm the task is still running.

Rate limiting - Reduce concurrent requests or use run_async() for batch workloads.