Skip to main content
Version: v1.0

aixplain.v1.modules.model.rlm

RLM (Recursive Language Model) module for aiXplain SDK v1.

__author__​

Copyright 2026 The aiXplain SDK authors

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.

Author: aiXplain Team Description: RLM (Recursive Language Model) — orchestrates long-context analysis via an iterative REPL sandbox. The orchestrator model plans and writes Python code to chunk and explore a large context; a worker model handles per-chunk analysis via llm_query() calls injected into the sandbox session.

RLM Objects​

class RLM(Model)

[view_source]

Recursive Language Model — long-context analysis with two execution modes.

RLM wraps two aiXplain models:

  • An orchestrator (powerful, expensive): plans and writes Python code to explore the context iteratively in a managed sandbox environment. Used by mode="recursive".
  • A worker (fast, cheap): called per chunk in both modes. In recursive mode it's invoked via llm_query() inside the sandbox; in parallel mode it's called directly, concurrently across chunks.

Three run modes are available via the mode argument to run():

  • "parallel" (cheap, fast, deterministic): chunk the context to a comfortable fraction of the worker's window (to mitigate context rot), call the worker in parallel on every chunk, then reduce the partial answers into a final answer. No orchestrator, no sandbox. Best for summarize/extract queries.
  • "rag" (cheapest at query time): chunk the context into small retrieval-grade pieces, upsert them into an aIR vector index, retrieve the top-k most relevant chunks for the query, then make a single worker call to synthesize an answer. The win is amortizing the upfront index build across many queries: set rag_index_id to reuse a pre-built index and skip create + upsert + delete on each call. Best for needle-in-haystack questions on very large contexts.
  • "recursive" (adaptive, expensive): the original iterative REPL loop where the orchestrator drives chunking and analysis. Best for multi-hop reasoning or queries that need to compare information across chunks.
  • "auto" (default): picks "recursive" if the query reads like multi-hop reasoning (e.g., contains words like "compare across", "inconsistencies", "verify against"); otherwise "parallel". "rag" is opt-in only — auto never picks it because it only beats parallel when the index is reused across many calls.

The recursive mode's sandbox is an aiXplain managed Python execution environment. Each recursive run() call gets its own isolated session (UUID), so variables persist across REPL iterations within a single run but are cleaned up afterwards.

Example usage:

from aixplain.factories import ModelFactory

rlm = ModelFactory.create_rlm(
orchestrator_model_id="<orchestrator-model-id>",
worker_model_id="<worker-model-id>",
)
response = rlm.run(data={
"context": very_long_document,
"query": "What are the key findings?",
})
print(response.data)
print(f"Completed in {response['iterations_used']} iterations.")

Attributes:

  • orchestrator Model - Root LLM that plans and writes REPL code.
  • worker Model - Sub-LLM used inside the sandbox via llm_query().
  • max_iterations int - Maximum orchestrator loop iterations before a forced final answer is requested.

__init__​

def __init__(id: Text,
name: Text = "RLM",
description:
Text = "Recursive Language Model for long-context analysis.",
orchestrator: Optional[Model] = None,
worker: Optional[Model] = None,
max_iterations: int = 10,
api_key: Optional[Text] = None,
supplier: Union[Dict, Text, Supplier, int] = "aiXplain",
rag_index_id: Optional[Text] = None,
rag_top_k: int = _RAG_DEFAULT_TOP_K,
rag_max_chunk_chars: int = _RAG_DEFAULT_MAX_CHUNK_CHARS,
**additional_info) -> None

[view_source]

Initialize a new RLM instance.

Arguments:

  • id Text - Identifier for this RLM instance.
  • name Text, optional - Display name. Defaults to "RLM".
  • description Text, optional - Description. Defaults to a generic string.
  • orchestrator Model, optional - Root LLM that drives the REPL loop. Must be set before calling run(). Defaults to None.
  • worker Model, optional - Sub-LLM called inside the sandbox via llm_query(). Must be set before calling run(). Defaults to None.
  • max_iterations int, optional - Maximum orchestrator iterations. Defaults to 10.
  • api_key Text, optional - API key. Defaults to config.TEAM_API_KEY.
  • supplier Union[Dict, Text, Supplier, int], optional - Supplier. Defaults to "aiXplain".
  • rag_index_id Text, optional - ID of a pre-built aIR index for mode="rag". When set, the index is reused (skipping create + upsert + delete on each run()). When None, an ephemeral index is created per RAG run and deleted afterwards. Defaults to None.
  • rag_top_k int, optional - Number of chunks retrieved from the aIR index per query in mode="rag". Defaults to 10.
  • rag_max_chunk_chars int, optional - Upper bound on a single RAG chunk (chars), to stay under the embedding model's input limit. Default ~30K chars is a generic middle ground that fits 8K-token models (ada-002, text-embedding-3, BGE-M3). Tune down for 512-token models (multilingual-E5, Jina CLIP) or up when the backing model supports it. Defaults to 30000.
  • **additional_info - Additional metadata stored on the instance.

run​

def run(data: Union[Text, Dict],
name: Text = "rlm_process",
timeout: float = 600,
parameters: Optional[Dict] = None,
wait_time: float = 0.5,
stream: bool = False,
mode: Text = "auto") -> ModelResponse

[view_source]

Run the RLM over a (potentially large) context, dispatching by mode.

mode selects the execution strategy:

  • "parallel": deterministic chunk + parallel worker calls + reduce. Cheap, fast, predictable. No orchestrator or sandbox needed. Best for summarize/extract queries. Chunks are sized to a comfortable fraction of the worker's window to mitigate context rot.
  • "rag": chunk + upsert to an aIR vector index + top-k retrieve + single worker synthesis. If self.rag_index_id is set, the index is reused (cheapest path); otherwise a temporary index is created and deleted per run. Best for needle-in-haystack queries on very large reusable contexts.
  • "recursive": the iterative REPL loop where the orchestrator drives chunking and analysis in a sandbox. Expensive but adaptive. Best for multi-hop reasoning that needs to compare information across chunks.
  • "auto" (default): picks "recursive" when the query reads like multi-hop reasoning (keywords such as "compare across", "inconsistencies", "verify against", "step by step"); otherwise "parallel". "rag" is opt-in only.

Arguments:

  • data Union[Text, Dict] - Input data. Accepted formats:

    • str (raw text): Treated directly as the context content; a default query is used.
    • str (HTTP/HTTPS URL): Content is downloaded automatically. .json URLs or application/json responses are parsed into a dict/list; all other content is decoded as plain text.
    • str (file path): If the string points to an existing file, the file is read automatically. .json files are parsed into a dict/list; all other text formats are read as a plain string.
    • pathlib.Path: Resolved and read exactly like a file-path string.
    • dict: Must contain "context" (required) and optionally "query" (defaults to a generic analysis prompt). The value of "context" itself may also be a URL, a file path, or a pathlib.Path — it will be resolved the same way.
  • name Text, optional - Identifier used in log messages. Defaults to "rlm_process".

  • timeout float, optional - Maximum total wall-clock time in seconds. Applies to "recursive" mode only. Defaults to 600.

  • parameters Optional[Dict], optional - Reserved for future use.

  • wait_time float, optional - Kept for API compatibility. Unused.

  • stream bool, optional - Unsupported. Must be False.

  • mode Text, optional - "auto", "parallel", or "recursive". Defaults to "auto".

Returns:

  • ModelResponse - Standard response with:

    • data: The final answer string.
    • completed: True on success.
    • run_time: Total elapsed seconds.
    • used_credits: Total credits consumed across all model calls.
    • iterations_used: Recursive mode → orchestrator iterations. Parallel mode → worker calls made in the map step (i.e. number of chunks), or 1 for the single-call fast path.

Raises:

  • AssertionError - If worker is not set, stream=True, or mode is invalid. recursive mode additionally requires orchestrator to be set.
  • ValueError - If data is a dict missing the "context" key, or an unsupported type.

run_async​

def run_async(data: Union[Text, Dict],
name: Text = "rlm_process",
parameters: Optional[Dict] = None) -> ModelResponse

[view_source]

Not supported for RLM.

Raises:

  • NotImplementedError - Always. Use run() instead.

run_stream​

def run_stream(data: Union[Text, Dict], parameters: Optional[Dict] = None)

[view_source]

Not supported for RLM.

Raises:

  • NotImplementedError - Always.

from_dict​

@classmethod
def from_dict(cls, data: Dict) -> "RLM"

[view_source]

Create an RLM instance from a dictionary representation.

Reconstructs the RLM from a dict produced by to_dict(). The orchestrator and worker models are fetched from the aiXplain platform using their stored IDs, so a valid API key and network access are required.

Arguments:

  • data Dict - Dictionary as produced by to_dict(), containing:
    • id: RLM instance identifier.
    • name: Display name.
    • description: Description string.
    • api_key: API key for authentication.
    • supplier: Supplier information.
    • orchestrator_model_id: ID of the orchestrator model.
    • worker_model_id: ID of the worker model.
    • max_iterations: Maximum orchestrator loop iterations.
    • additional_info: Extra metadata (optional).

Returns:

  • RLM - A fully configured RLM instance with orchestrator and worker models loaded, ready to call run().

Raises:

  • Exception - If either model ID cannot be fetched from the platform.

to_dict​

def to_dict() -> Dict

[view_source]

Convert the RLM instance to a dictionary representation.

Extends the base Model.to_dict() with RLM-specific fields: orchestrator model ID, worker model ID, and max_iterations. The orchestrator and worker are stored as their model IDs (not full objects) so the dict is JSON-serializable and can be used to reconstruct the instance via ModelFactory.create_rlm().

Returns:

  • Dict - A dictionary containing all base model fields plus:
    • orchestrator_model_id: ID of the orchestrator model.
    • worker_model_id: ID of the worker model.
    • max_iterations: Maximum orchestrator loop iterations.

__repr__​

def __repr__() -> str

[view_source]

Return a string representation of this RLM instance.