Run a Function
A published Function can be run two ways:- Real time. One input, one synchronous answer. Use it for single requests and low-QPS integrations, and for Functions with image or PDF inputs.
- Batch. A job over a table, dataframe, or file. Use it for larger datasets, from thousands to millions of rows. Note that it is text-only.
Install and authenticate
sutro login.
Run one input
Request
POST /v1/functions/{name}/run takes exactly
one body field, input. Its keys must match the Function’s configured input
fields; every field is required, and unknown fields are rejected. Values that
are not strings are converted the same way Batch converts a row: null becomes
an empty string, objects and lists become JSON, and anything else becomes its
plain string form. A bare string is accepted in place of the object when the
Function has exactly one text input.
Sampling parameters cannot be overridden per request. The Function’s published
settings are what make its answers reproducible.
Response
Confidence scores
confidence reports how settled the answer is, using the same technique as
the Batch API. Answers at or below 0.6 are the ones worth a human look.
Treat the score as a routing signal rather than a probability: send
low-confidence answers to a review step, a fallback path, or a person, and let
the rest through. The inputs behind them are also the most useful ones to label
in the next iteration of the Function.
Scoring adds cost and latency to a request. If you will not act on the score,
turn it off per request with confidence_scoring=False in the SDK or
"confidence_scoring": false in the request body. confidence is then null.
The answer still appears in Monitoring but is never added to
Needs Review.
Image and PDF inputs
Functions with an image or PDF input field are real-time only, since Batch is currently text-only. An asset field takes one of three values.
PNG, JPEG, WebP, and PDF are accepted, up to 20 MB per asset. The bytes are
sniffed, so a declared
type or mime_type that disagrees with the actual
content is rejected rather than passed to the model.
Limits
A rate-limited request returns
429 with a Retry-After header; wait that many
seconds before retrying. The SDK surfaces it as
SutroRateLimitError.retry_after. Deployment operators can change these
defaults with the HARMONIZE_REALTIME_* environment variables.
Errors
Failures return{"detail": ..., "code": ..., "request_id": ...}.
Requests are never retried automatically, in the SDK or the API: a Function run
costs tokens, and a caller waiting on an answer should decide for itself
whether to try again.
Submit many rows instead
For anything larger than a few hundred inputs, Batch is the better developer experience. It takes a list of dictionaries, Pandas or Polars DataFrames, local CSV or Parquet files, and HTTPS download URLs. It has no request rate limits and comfortably handles hundreds of thousands to millions of rows.batch_run_function() and
Run a Function in Batch.