aixplain.v1.modules.model.rlm
RLM (Recursive Language Model) module for aiXplain SDK v1.
__author__
Copyright 2026 The aiXplain SDK authors
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.
Author: aiXplain Team Description: RLM (Recursive Language Model) — orchestrates long-context analysis via an iterative REPL sandbox. The orchestrator model plans and writes Python code to chunk and explore a large context; a worker model handles per-chunk analysis via llm_query() calls injected into the sandbox session.
RLM Objects
class RLM(Model)
Recursive Language Model — long-context analysis with two execution modes.
RLM wraps two aiXplain models:
- An orchestrator (powerful, expensive): plans and writes Python code to
explore the context iteratively in a managed sandbox environment.
Used by
mode="recursive". - A worker (fast, cheap): called per chunk in both modes. In recursive
mode it's invoked via
llm_query()inside the sandbox; in parallel mode it's called directly, concurrently across chunks.
Three run modes are available via the mode argument to run():
"parallel"(cheap, fast, deterministic): chunk the context to a comfortable fraction of the worker's window (to mitigate context rot), call the worker in parallel on every chunk, then reduce the partial answers into a final answer. No orchestrator, no sandbox. Best for summarize/extract queries."rag"(cheapest at query time): chunk the context into small retrieval-grade pieces, upsert them into an aIR vector index, retrieve the top-k most relevant chunks for the query, then make a single worker call to synthesize an answer. The win is amortizing the upfront index build across many queries: setrag_index_idto reuse a pre-built index and skip create + upsert + delete on each call. Best for needle-in-haystack questions on very large contexts."recursive"(adaptive, expensive): the original iterative REPL loop where the orchestrator drives chunking and analysis. Best for multi-hop reasoning or queries that need to compare information across chunks."auto"(default): picks"recursive"if the query reads like multi-hop reasoning (e.g., contains words like "compare across", "inconsistencies", "verify against"); otherwise"parallel"."rag"is opt-in only — auto never picks it because it only beatsparallelwhen the index is reused across many calls.
The recursive mode's sandbox is an aiXplain managed Python execution
environment. Each recursive run() call gets its own isolated session
(UUID), so variables persist across REPL iterations within a single run
but are cleaned up afterwards.
Example usage:
from aixplain.factories import ModelFactory
rlm = ModelFactory.create_rlm(
orchestrator_model_id="<orchestrator-model-id>",
worker_model_id="<worker-model-id>",
)
response = rlm.run(data={
"context": very_long_document,
"query": "What are the key findings?",
})
print(response.data)
print(f"Completed in {response['iterations_used']} iterations.")
Attributes:
orchestratorModel - Root LLM that plans and writes REPL code.workerModel - Sub-LLM used inside the sandbox viallm_query().max_iterationsint - Maximum orchestrator loop iterations before a forced final answer is requested.
__init__
def __init__(id: Text,
name: Text = "RLM",
description:
Text = "Recursive Language Model for long-context analysis.",
orchestrator: Optional[Model] = None,
worker: Optional[Model] = None,
max_iterations: int = 10,
api_key: Optional[Text] = None,
supplier: Union[Dict, Text, Supplier, int] = "aiXplain",
rag_index_id: Optional[Text] = None,
rag_top_k: int = _RAG_DEFAULT_TOP_K,
rag_max_chunk_chars: int = _RAG_DEFAULT_MAX_CHUNK_CHARS,
**additional_info) -> None
Initialize a new RLM instance.
Arguments:
idText - Identifier for this RLM instance.nameText, optional - Display name. Defaults to "RLM".descriptionText, optional - Description. Defaults to a generic string.orchestratorModel, optional - Root LLM that drives the REPL loop. Must be set before callingrun(). Defaults to None.workerModel, optional - Sub-LLM called inside the sandbox viallm_query(). Must be set before callingrun(). Defaults to None.max_iterationsint, optional - Maximum orchestrator iterations. Defaults to 10.api_keyText, optional - API key. Defaults toconfig.TEAM_API_KEY.supplierUnion[Dict, Text, Supplier, int], optional - Supplier. Defaults to "aiXplain".rag_index_idText, optional - ID of a pre-built aIR index formode="rag". When set, the index is reused (skipping create + upsert + delete on eachrun()). WhenNone, an ephemeral index is created per RAG run and deleted afterwards. Defaults toNone.rag_top_kint, optional - Number of chunks retrieved from the aIR index per query inmode="rag". Defaults to 10.rag_max_chunk_charsint, optional - Upper bound on a single RAG chunk (chars), to stay under the embedding model's input limit. Default ~30K chars is a generic middle ground that fits 8K-token models (ada-002, text-embedding-3, BGE-M3). Tune down for 512-token models (multilingual-E5, Jina CLIP) or up when the backing model supports it. Defaults to 30000.**additional_info- Additional metadata stored on the instance.
run
def run(data: Union[Text, Dict],
name: Text = "rlm_process",
timeout: float = 600,
parameters: Optional[Dict] = None,
wait_time: float = 0.5,
stream: bool = False,
mode: Text = "auto") -> ModelResponse
Run the RLM over a (potentially large) context, dispatching by mode.
mode selects the execution strategy:
"parallel": deterministic chunk + parallel worker calls + reduce. Cheap, fast, predictable. No orchestrator or sandbox needed. Best for summarize/extract queries. Chunks are sized to a comfortable fraction of the worker's window to mitigate context rot."rag": chunk + upsert to an aIR vector index + top-k retrieve + single worker synthesis. Ifself.rag_index_idis set, the index is reused (cheapest path); otherwise a temporary index is created and deleted per run. Best for needle-in-haystack queries on very large reusable contexts."recursive": the iterative REPL loop where the orchestrator drives chunking and analysis in a sandbox. Expensive but adaptive. Best for multi-hop reasoning that needs to compare information across chunks."auto"(default): picks"recursive"when the query reads like multi-hop reasoning (keywords such as "compare across", "inconsistencies", "verify against", "step by step"); otherwise"parallel"."rag"is opt-in only.
Arguments:
-
dataUnion[Text, Dict] - Input data. Accepted formats:str(raw text): Treated directly as the context content; a default query is used.str(HTTP/HTTPS URL): Content is downloaded automatically..jsonURLs orapplication/jsonresponses are parsed into a dict/list; all other content is decoded as plain text.str(file path): If the string points to an existing file, the file is read automatically..jsonfiles are parsed into a dict/list; all other text formats are read as a plain string.pathlib.Path: Resolved and read exactly like a file-path string.dict: Must contain"context"(required) and optionally"query"(defaults to a generic analysis prompt). The value of"context"itself may also be a URL, a file path, or apathlib.Path— it will be resolved the same way.
-
nameText, optional - Identifier used in log messages. Defaults to"rlm_process". -
timeoutfloat, optional - Maximum total wall-clock time in seconds. Applies to"recursive"mode only. Defaults to 600. -
parametersOptional[Dict], optional - Reserved for future use. -
wait_timefloat, optional - Kept for API compatibility. Unused. -
streambool, optional - Unsupported. Must be False. -
modeText, optional -"auto","parallel", or"recursive". Defaults to"auto".
Returns:
-
ModelResponse- Standard response with:data: The final answer string.completed: True on success.run_time: Total elapsed seconds.used_credits: Total credits consumed across all model calls.iterations_used: Recursive mode → orchestrator iterations. Parallel mode → worker calls made in the map step (i.e. number of chunks), or 1 for the single-call fast path.
Raises:
AssertionError- Ifworkeris not set,stream=True, ormodeis invalid.recursivemode additionally requiresorchestratorto be set.ValueError- Ifdatais a dict missing the"context"key, or an unsupported type.
run_async
def run_async(data: Union[Text, Dict],
name: Text = "rlm_process",
parameters: Optional[Dict] = None) -> ModelResponse
Not supported for RLM.
Raises:
NotImplementedError- Always. Userun()instead.
run_stream
def run_stream(data: Union[Text, Dict], parameters: Optional[Dict] = None)
Not supported for RLM.
Raises:
NotImplementedError- Always.
from_dict
@classmethod
def from_dict(cls, data: Dict) -> "RLM"
Create an RLM instance from a dictionary representation.
Reconstructs the RLM from a dict produced by to_dict(). The
orchestrator and worker models are fetched from the aiXplain platform
using their stored IDs, so a valid API key and network access are
required.
Arguments:
dataDict - Dictionary as produced byto_dict(), containing:- id: RLM instance identifier.
- name: Display name.
- description: Description string.
- api_key: API key for authentication.
- supplier: Supplier information.
- orchestrator_model_id: ID of the orchestrator model.
- worker_model_id: ID of the worker model.
- max_iterations: Maximum orchestrator loop iterations.
- additional_info: Extra metadata (optional).
Returns:
RLM- A fully configured RLM instance with orchestrator and worker models loaded, ready to callrun().
Raises:
Exception- If either model ID cannot be fetched from the platform.
to_dict
def to_dict() -> Dict
Convert the RLM instance to a dictionary representation.
Extends the base Model.to_dict() with RLM-specific fields:
orchestrator model ID, worker model ID, and max_iterations.
The orchestrator and worker are stored as their model IDs (not full
objects) so the dict is JSON-serializable and can be used to
reconstruct the instance via ModelFactory.create_rlm().
Returns:
Dict- A dictionary containing all base model fields plus:- orchestrator_model_id: ID of the orchestrator model.
- worker_model_id: ID of the worker model.
- max_iterations: Maximum orchestrator loop iterations.
__repr__
def __repr__() -> str
Return a string representation of this RLM instance.