Skip to main content

Run a Function

A published Function can be run two ways:
  • Real time. One input, one synchronous answer. Use it for single requests and low-QPS integrations, and for Functions with image or PDF inputs.
  • Batch. A job over a table, dataframe, or file. Use it for larger datasets, from thousands to millions of rows. Note that it is text-only.
Both run the same published prompt, model, output schema, and sampling defaults, so the answers agree. Open Integrate on a ready Function to see code snippets to integrate into your pipeline or script.

Install and authenticate

Use the origin of the deployment where you published the Function and a deployment API key. Verify your connection before submitting a job. You can persist the same values with sutro login.

Run one input

The same call over HTTP:
Or from the command line:

Request

POST /v1/functions/{name}/run takes exactly one body field, input. Its keys must match the Function’s configured input fields; every field is required, and unknown fields are rejected. Values that are not strings are converted the same way Batch converts a row: null becomes an empty string, objects and lists become JSON, and anything else becomes its plain string form. A bare string is accepted in place of the object when the Function has exactly one text input. Sampling parameters cannot be overridden per request. The Function’s published settings are what make its answers reproducible.

Response

Confidence scores

confidence reports how settled the answer is, using the same technique as the Batch API. Answers at or below 0.6 are the ones worth a human look. Treat the score as a routing signal rather than a probability: send low-confidence answers to a review step, a fallback path, or a person, and let the rest through. The inputs behind them are also the most useful ones to label in the next iteration of the Function. Scoring adds cost and latency to a request. If you will not act on the score, turn it off per request with confidence_scoring=False in the SDK or "confidence_scoring": false in the request body. confidence is then null. The answer still appears in Monitoring but is never added to Needs Review.

Image and PDF inputs

Functions with an image or PDF input field are real-time only, since Batch is currently text-only. An asset field takes one of three values.
PNG, JPEG, WebP, and PDF are accepted, up to 20 MB per asset. The bytes are sniffed, so a declared type or mime_type that disagrees with the actual content is rejected rather than passed to the model.

Limits

A rate-limited request returns 429 with a Retry-After header; wait that many seconds before retrying. The SDK surfaces it as SutroRateLimitError.retry_after. Deployment operators can change these defaults with the HARMONIZE_REALTIME_* environment variables.

Errors

Failures return {"detail": ..., "code": ..., "request_id": ...}. Requests are never retried automatically, in the SDK or the API: a Function run costs tokens, and a caller waiting on an answer should decide for itself whether to try again.

Submit many rows instead

For anything larger than a few hundred inputs, Batch is the better developer experience. It takes a list of dictionaries, Pandas or Polars DataFrames, local CSV or Parquet files, and HTTPS download URLs. It has no request rate limits and comfortably handles hundreds of thousands to millions of rows.
See batch_run_function() and Run a Function in Batch.

Run with your own provider

The OpenAI-compatible endpoints section in Integrate provides the optimized system prompt, output schema, and recommended model. Copy these artifacts into your own compatible provider when you want to run the Function’s configuration on your own infrastructure. That path runs the exported configuration: it does not invoke the published Function by name, and it does not return a confidence score.