Local SDK preview

From requests to results.

Prepare a JSONL file, create a batch, and retrieve the output using Python.

This guide describes the current local SDK. Pilot setup is arranged with Firmbatch; there is no public package installation or hosted API key flow documented here.

Before you begin

You need access to the Firmbatch code, a Python environment prepared for the pilot, and an operator-configured execution profile. That profile defines the available models and cloud resources. Run these examples where from firmbatch import LocalClient can be imported.

LocalClient() loads the default local profile, or the profile named by FIRMBATCH_PROFILE. Execution can incur cloud charges under that configuration. The coordinating machine must remain on while work runs.

Discuss your first pilot to arrange setup and agree on the model, processing location, and workload.

1. Prepare requests

Save one JSON object per line in requests.jsonl. Each request needs a unique custom_id and a body containing chat messages.

requests.jsonl
{"custom_id":"review-001","body":{"messages":[{"role":"user","content":"Classify as positive, negative, or neutral. Reply with one word: Support took weeks and never solved the problem."}],"temperature":0,"max_tokens":8}}
{"custom_id":"review-002","body":{"messages":[{"role":"user","content":"Classify as positive, negative, or neutral. Reply with one word: The team was helpful and resolved my issue."}],"temperature":0,"max_tokens":8}}

The body uses an OpenAI-style chat-completion shape. This does not imply compatibility with the hosted OpenAI Batch API. The model is chosen when creating the batch; a model field inside the body does not select it.

FieldRule
custom_idRequired and unique. 1–128 ASCII characters: A–Z, a–z, 0–9, periods, underscores, colons, or hyphens; start with an ASCII letter or digit.
body.messagesRequired, non-empty list. Every message needs a string role and a content field.
methodOptional. If supplied, it must be POST.
urlOptional. If supplied, it must be /v1/chat/completions.

2. Run a batch

The example model name is an alias in the example operator profile. Check client.models for the aliases actually configured for your pilot.

run_batch.py
from firmbatch import LocalClient

client = LocalClient()
print("Configured models:", client.models)

batch = client.batches.create(
    model="llama-3.1-8b-instruct",
    input_file="requests.jsonl",
    name="review-classification",
)
print("Batch ID:", batch.id)
batch.wait()

print("Status:", batch.status)
if batch.error:
    print("Batch error:", batch.error)
if batch.output_file is not None:
    batch.download("results.jsonl")

For this small-file workflow, create() validates and stores the input. wait() drives execution and prints progress. Keep that process running. Finishing the wait does not by itself mean every request succeeded: inspect the batch status and output records.

3. Track progress

Keep the batch ID to retrieve it from the same local profile and state directory. Status and progress are read from durable local state.

Inspect an existing batch
batch = client.batches.retrieve("YOUR_BATCH_ID")
print(batch.status)
print(batch.progress.summary())
print(batch.error)

# Resume or wait for the batch using its existing state.
batch.wait()

A batch can wait for capacity, start its model, process requests, and finalize results. Its terminal status is completed, failed, or cancelled. A capacity wait is not a guaranteed completion time.

Use batch.cancel() to request cancellation. Completed responses are kept. The small-file and manifest workflows handle an interrupted wait differently; see the larger-dataset section before leaving a job unattended.

4. Read results

For a small-file batch, download() saves a JSONL file and results() returns the parsed records. Match results to inputs using custom_id.

Handle each request outcome
if batch.output_file is not None:
    for result in batch.results():
        if result["status"] == "completed":
            response = result["response"]["body"]
            print(result["custom_id"], response)
        else:
            print(result["custom_id"], result["status"], result["error"])

Each output record contains custom_id, status, response, and error. A completed response includes status_code: 200 and the model response in body. Other record statuses are failed, cancelled, and unfinished.

A completed request means inference returned successfully. It does not certify that its answer meets your quality criteria; evaluate a representative sample before scaling up.

Larger datasets

The manifest workflow accepts pre-sharded JSONL in object storage. An operator prepares and validates the manifest, including checksums and unique request IDs. Supply input_manifest instead of input_file.

Operator-prepared manifest
batch = client.batches.create(
    model="llama-3.1-8b-instruct",
    input_manifest="s3://YOUR_BUCKET/YOUR_PREFIX/manifest.json",
)
batch.wait()
print(batch.status, batch.error)
batch.download("result-shards")

for result in batch.iter_results():
    print(result["custom_id"], result["status"])

Here create() starts a background coordinator on the local machine and returns. wait() only watches; interrupting it leaves the batch running. Use batch.cancel() to stop it. The local machine still needs to stay on.

For manifest batches, download() takes a directory and writes result shards. iter_results() reads one shard at a time; results() is not available. Partial results may exist even when a batch fails or is cancelled.

SDK reference

Call or propertyPurpose
client.modelsList model aliases in the current operator profile.
client.files.upload(path)Validate and store a small JSONL file. Pass the returned file to input_file.
client.batches.create(...)Create a batch with exactly one of input_file or input_manifest.
client.batches.retrieve(id)Open a batch in the same profile's local state.
client.batches.list()List batches recorded in that state.
batch.wait()Drive a small-file batch, or watch a manifest batch.
batch.status, batch.errorRead the batch state and failure explanation, when present.
batch.download(destination)Save a result file, or a directory of shards for a manifest batch.
batch.cancel()Request cancellation while keeping completed output.

Limits and errors

The small-file parser accepts UTF-8 JSONL up to 200 MiB, at most 50,000 requests, and at most 1 MiB per line. Use the operator-prepared manifest workflow for larger input.

Streaming responses and multiple completions per request are not supported. Omit stream, stream_options, and n. Small-file records accept only the four envelope fields listed above; optional metadata belongs to the manifest workflow.

Invalid input raises InputValidationError. An unconfigured model raises UnknownModelError. Reading a small-file result before output exists raises OutputNotAvailable. Execution failures are also reported through batch.status and batch.error.

This preview documents the local implementation as reviewed on 8 October 2026. Model availability, deployment locations, pricing, and service terms are agreed separately for each pilot.

Discuss a pilot