Skip to main content
Evaluators in LangSmith are workspace-level resources. You can attach a single evaluator to multiple tracing projects and datasets, so you can apply consistent evaluation logic across your work without recreating it each time.
Evaluator scores are a high-priority signal for the LangSmith Engine: it pulls low-scoring traces when choosing what to analyze, so attaching an evaluator to a project sharpens the issues Engine finds there.

View evaluators

In the LangSmith UI, select Evaluators in the left sidebar to view all evaluators in your workspace. The evaluators table shows the following columns:

Create an evaluator

You can create an evaluator in the LangSmith UI or programmatically with the SDK. Evaluators created either way are workspace-level resources that appear in the Evaluators table.

Create an evaluator in the UI

  1. In the LangSmith UI, select Evaluators in the left sidebar.
  2. Click + Evaluator to open the new evaluator panel.
  3. The panel lets you:
    • Create from scratch: Build a new LLM-as-a-Judge or Code evaluator.
    • Add a LangChain Tuned Evaluator: Attach a specialized judge managed by LangChain to a compatible tracing project without configuring a prompt, model, or API key.
    • Create from a template: Start from a ready-made evaluator (also known as a prebuilt evaluator) for common evaluation patterns. A Recommended section surfaces popular templates first, followed by templates organized by the following categories:
You can also add an evaluator directly from a tracing project or dataset. In that flow, you can additionally attach an existing evaluator from your workspace, or create a Composite evaluator. Refer to Set up LLM-as-a-judge online evaluators and Automatically run evaluators on experiments.

Create an evaluator with the SDK

Use the LangSmith SDK to create evaluators programmatically. The SDK is available for Python and TypeScript. Evaluators created through the SDK appear in the Evaluators table alongside those created in the UI.
Managing evaluators through the SDK requires langsmith>=0.9.8 (Python, PyPI) or langsmith>=0.7.16 (TypeScript, npm).
To create an LLM-as-a-judge evaluator and to retrieve, update, list, or delete evaluators, refer to Manage evaluators with the SDK.

View evaluator details

Click any evaluator in the table to open its detail view. The detail view has four tabs:
  • Overview: The evaluator’s feedback configuration and prompt or code definition.
  • Traces: Traces processed by this evaluator across all attached resources.
  • Logs: Execution logs for this evaluator across all attached resources.
  • Projects & Datasets: The tracing projects and datasets this evaluator is attached to, with each attachment’s weekly spend and limit.

Edit an evaluator

Open an evaluator. In the Overview tab, click the Edit evaluator icon to open the Configure Evaluator panel. Update the evaluator’s configuration. Click Save. Because the evaluator is shared, changes apply across all tracing projects and datasets it is attached to.

Manage evaluator trace retention

When an online evaluator scores a trace, it attaches feedback to the trace. This can auto-upgrade the trace to extended retention, depending on the evaluator’s retention setting. Extended retention keeps the trace longer but costs more. When you set up an online evaluator on a tracing project, you can opt out of this upgrade so that scored traces stay at the project’s base retention. This control is available only when the project’s default retention is the base tier. If the project defaults to extended retention (set at the project or workspace level), traces scored by the evaluator follow that default and the option is locked. To opt out of extending retention for scored traces:
  1. When you create or edit an online evaluator, set the source to a tracing project, rather than a dataset.
  2. Expand the Advanced section in the evaluator configuration panel.
  3. Clear Extend trace retention.
The change applies to traces scored after you save the evaluator. Existing scored traces keep their current retention tier. The Extend trace retention toggle described above applies to both trace-level and thread-level (multi-turn) online evaluators. For more information on multi-turn evaluators, see Set up multi-turn online evaluators.

Include extended stats

Use Include extended stats (feedback, costs, tokens) in a run-level evaluator to evaluate feedback statistics, token usage, or cost data from the run. The feedback_stats field contains feedback statistics, including the number and average for each feedback key. This option is not available for multi-turn (thread-level) evaluators. LangSmith fetches additional data for evaluators with this option enabled. Enable it only when your evaluation logic or prompt requires these fields. To include extended stats:
  1. When you create or edit a run-level evaluator, expand the Advanced section in the evaluator configuration panel.
  2. Select Include extended stats (feedback, costs, tokens).

Access extended stats

Code evaluators receive these fields in the run object, including feedback_stats, total_tokens, and total_cost. They also receive individual feedback records in run["feedback"], regardless of whether you enable extended stats. For LLM-as-a-judge evaluators, map the corresponding run.* field to a prompt variable. For example, map {{correctness_average}} to run.feedback_stats.correctness.avg to include a correctness feedback average in the prompt. The available run.* fields also include prompt and completion token, cost, and detail fields. For the full run data schema, see Run data format.

Chain evaluators

Extended stats can be used to chain evaluators, where a code evaluator reads a score that another evaluator already produced. Filter the second evaluator on the first evaluator’s feedback key (for example, has(feedback_key, "answer_usefulness")), then enable extended stats. The filter makes the code evaluator run only after the score exists, and extended stats make the score readable at run["feedback_stats"]["answer_usefulness"]["avg"]. The filter matches on the feedback key rather than the evaluator that produced it, so feedback from any source with that key triggers the code evaluator. To configure the filter, see Apply a filter to runs that trigger the evaluator.

Delete an evaluator

You cannot delete an evaluator while it is attached to a tracing project or dataset. To delete an evaluator:
  1. In the LangSmith UI, select Evaluators in the left sidebar.
  2. Select the evaluator you want to delete.
  3. Open the Projects & Datasets tab. For each attached tracing project and dataset, select Detach in the Actions menu at the right of the row.
  4. Return to the Evaluators page and click Delete at the top of the page.