- Set up your harness — connect your existing agent to Raindrop.
- Create an eval — describe what a good reply looks like.
- Configure your replay suite — choose the agent, inputs, and evaluators.
- Run — generate replies and review the scores.
- Use the playground — try changes on a few rows, then run the full suite.
1. Set up your harness
A harness is a small adapter around your agent that lets Raindrop call it with test inputs. Your agent still runs on your server.What the harness does
One URL, such ashttps://your-app.com/api/evals, handles two jobs:
- Discover your agents: Raindrop sends a
GETrequest to learn which agents and parameters to show in the form. This does not run your agent. - Run your agent: Raindrop sends a
POSTrequest when you start a test. The SDK loads the selected dataset rows, calls your agent for each row, and sends the traces back to Raindrop.
Install the SDK and prepare your keys
defineAgent and createEvalHandler. If your installed version does not include them, use the replay-enabled release provided by your Raindrop team before continuing.
Add these variables to your server environment:
RAINDROP_API_KEYlets the SDK read the run and upload its traces to Raindrop. Use an organization API key, not a telemetry write key. The key must have access to the project you are testing.EVAL_ENDPOINT_SECRETprotects your endpoint. You will enter this same secret in Raindrop’s endpoint configuration so Raindrop can call your server.
Wrap your agent
Here is a complete Next.js App Router route. Put it inapp/api/evals/route.ts, then replace the answerCustomer import and call with your own agent function.
The example assumes your agent accepts a message and options, then returns its final reply as a string.
What each wrapper does
defineAgent describes the agent. Its name appears in the Agent dropdown. Its slug is the stable identifier used to call it. Its run function receives one dataset row at a time.
withSpan records your agent call. It captures the input and the returned reply. Keep the agent’s work inside the callback and await it. If your agent returns a response object, return the actual final text from that object. If it streams, collect and return the complete reply.
If your agent already uses Raindrop tracing, keep that instrumentation. You do not need to add a second copy of its existing root wrapper; the replay handler associates work performed inside run with the right dataset row. Existing model and tool instrumentation adds detail to the trace.
createEvalHandler supplies the HTTP handler. It handles discovery, validates parameters, claims the run, runs the rows, and uploads the results. Keep both GET and POST behind the same authentication.
Using a different framework? Mount handleRequest on a route that accepts a standard Web Request and returns a Response, or use your framework’s adapter. The agent definition stays the same.
Check agent discovery
Start your application and call the route:agents array with your support agent, its name, and the prompt and temperature fields. A current handler also returns concurrencySupported: true.
This is the same discovery request Raindrop makes. If this check fails, fix it before starting a test:
- 401: the endpoint secret is missing or does not match.
- 404 or HTML: the URL is not pointing to the replay route.
- An empty agent list: check the
agentsarray passed tocreateEvalHandler.
POST. Start a test in Raindrop; it creates the run before calling your endpoint.
Your harness is ready when discovery returns your agent and its parameters. Deploy it to a server reachable by Raindrop over HTTPS. For local development, use an authenticated HTTPS tunnel; hosted Raindrop cannot reach localhost on your laptop.
2. Create an eval
An eval checks one aspect of your agent’s reply. Start with a single clear question, such as “Did the agent answer the user’s question without unnecessary detail?” In Evals, create an evaluator and describe what it should check in plain language. For example:Check whether the reply is concise and helpful. It should answer the user’s question directly, avoid repeating itself, and include only the detail needed to take the next step.Review the generated policy and adjust it to match your expectations. If you have examples or human-labelled comparisons, attach them with Add reference. For a judge built from human labels, use Calibrate to check its agreement with those labels before relying on its scores. Choose the kind of check that matches your goal:
- Score one reply: assess helpfulness, correctness, or another quality independently.
- Compare two replies: use a pairwise evaluator to decide whether a new reply is better than a reference.
3. Configure your replay suite
A replay suite brings together your agent, an input dataset, and the evals that will check its replies. A run is one execution of that setup. If you do not have an input dataset yet, create or upload one in Datasets. Start with around 10 representative inputs. Inputs alone are enough to generate replies; you do not need to write example answers just to get started. Open Evals and the Test your agent form.- Keep Run in → Custom endpoint selected.
- Under Agent endpoint, choose New endpoint.
- Give it a recognizable name, such as Support staging.
- Enter the full route URL:
https://your-app.com/api/evals. - Add a header with the name
Authorizationand the valueBearer YOUR_ENDPOINT_SECRET. ReplaceYOUR_ENDPOINT_SECRETwith the value ofEVAL_ENDPOINT_SECRETon your server. IncludeBearerand the space. - If the Concurrency field is available, set it to the number of dataset rows your harness can handle at once. Save the endpoint.
- Choose Support agent from the discovered agents.
- Choose your Input dataset. Start with around 10 rows so you can check the setup quickly.
- Expand Parameters if you want to change the prompt or temperature.
- Select the Evals you created and check the name. It starts with your dataset name, and you can edit it.
4. Run
Click Run suite. You will land on the run, where you can follow each input through execution and grading. Raindrop fixes the dataset version and run configuration, then calls your endpoint. The SDK runs your agent on those inputs and uploads the traces. The selected evaluators grade the resulting outputs separately. Execution completed means your agent finished. Grading may still be running. Your endpoint does not need to call the evaluators itself. Concurrency controls how many dataset rows your agent handles at once, not how many judges run. Start low if your agent uses shared state or has model rate limits. Open individual results to inspect the input, generated reply, and evaluator outcome. Check a few examples yourself before drawing conclusions from the overall scores. When grading finishes, you have a recorded run to return to and compare with future runs.5. Use the playground
Open the suite’s Playground when you want to try a change quickly.- Select a few dataset rows, or enter a sample percentage and click Sample.
- Expand Parameters and change the prompt or another setting your harness exposes.
- For a pairwise comparison, select a baseline run when you want to compare against its saved outputs. Without a selected baseline or a dataset reference, the pairwise cells show No baseline selected; you can still generate replies and run the other evals.
- Click Run selected and review the outputs and scores for those rows.
- Adjust and repeat. Use the same rows when comparing two prompt changes so you can see what actually improved.
- When you are happy with the change, click Run suite to run the full input dataset with your current configuration and record it in suite history.
Optional harness features
The five steps above are the core flow. These options help when you want to expose more controls or support more agents.Add more parameters
Theparameters declaration is also the form definition: Raindrop discovers it and renders the controls automatically.
You can declare text, number, boolean, and choice fields with z.string(), z.number(), z.boolean(), and z.enum(). Use .default() for sensible starting values and .describe() for a short explanation. These examples use Zod 3.
Parameter values arrive in run as parameters. They only affect your agent if you pass them into your agent function, as the example does with systemPrompt and temperature.
You can edit values for a test without saving a reusable preset. Use Save when you want to reuse a named configuration later. This makes it easy to try one prompt, change it, and run the same inputs again.
Register more than one agent
Register them on the same endpoint:defineAgent(...) result with a unique slug. Raindrop discovers all of them and shows their names in the Agent dropdown. Each agent can expose different parameters.
Share setup across rows
Usesetup and cleanup when a test needs a sandbox, temporary workspace, or other shared resource:
createTestWorkspace and runTracedSupportAgent stand for your application’s setup and traced agent functions. Setup happens once when execution reaches the first row. Rows share the returned environment, and cleanup runs after they settle, including after row failures. Keep each row’s conversation state separate. Cleanup cannot run after a process is killed, so temporary resources should also expire on their own.
If something does not work
Discovery works, but starting the run returns 401
Check which request failed. A 401 from your route means its Authorization header is wrong. A credential error when the SDK calls Raindrop meansRAINDROP_API_KEY is invalid, expired, or cannot access the selected project. A successful discovery request does not verify the SDK’s Raindrop credentials.
The run has no usable output
Return the final reply from your agent callback and await all the work that produces it. Confirm tracing is enabled and your traced call includes the input and output. Returning early while work continues in the background can leave the run without a complete trace.The request times out
The handler keeps the request open until agent execution finishes. Run it on infrastructure whose request timeout can cover the whole batch, including any proxy in front of it. A short-lived serverless route may be unsuitable for larger datasets. Do not return an early success response or puthandleEval into an unawaited background task.
Start with fewer rows, check the slowest agent call, and adjust concurrency within your agent’s limits. Raindrop’s current endpoint execution request timeout is 15 minutes.
A parameter is rejected
Check its declared type and limits. For example,temperature must be a number within the schema’s range. The SDK rejects invalid parameters before running the agent. Custom parameter values must be strings, finite numbers, or booleans; objects, arrays, and null are not supported.
The agent list does not show a new agent
Deploy the code that registers the agent, then checkGET on the exact endpoint URL saved in Raindrop. If the returned list is correct, select that endpoint again to refresh discovery.