Restructure evaluation docs

This commit is contained in:
Dominik Jain committed 2026-04-13 16:42:57 +02:00
1 parent c3a5652fc7
commit 46be9c4bd7
12 files changed
+260 -199

No files matched your search

@@ -0,0 +1,2 @@
# Evaluation
@@ -1,46 +1,26 @@
# Evaluation # Methodology
In this section we describe how we evaluate the performance of Serena's tools. In this section we describe the methodology we applied in evaluating the performance of Serena's tools.
The evaluation measures the **concrete delta** that Serena's tools provide on top of an agent's The evaluation measures the **concrete delta** that Serena's tools provide on top of an agent's
built-in capabilities (file reads, text edits, grep, shell, etc.). built-in capabilities (file reads, text edits, grep, shell, etc.).
Rather than a simple thumbs-up/thumbs-down, it produces a detailed, evidence-based report Rather than a simple thumbs-up/thumbs-down, it produces a detailed, evidence-based report
covering capabilities, efficiency and reliability. covering capabilities, efficiency and reliability.
## Why Not Benchmarks?
Standard coding benchmarks (SWE-bench, HumanEval, etc.) measure an agent's ability to solve
predefined tasks with a known correct answer. They are valuable for comparing models and agents,
but they are a poor fit for evaluating a **tool augmentation layer** like Serena for several reasons:
- **Benchmarks don't reflect real usage patterns.** Benchmark tasks are typically small, self-contained
problems that can be solved by reading and editing a handful of files. They rarely exercise the
workflows where Serena's tools shine — cross-file refactoring, navigating large codebases by symbol
structure, chaining multiple edits with stable addressing, or querying type hierarchies and
external dependencies. A benchmark score would mostly measure performance on tasks where Serena
is not expected to help.
- **Results would not generalise to the user's project.** Serena's value depends on the codebase
(size, language, complexity), the agent (model, built-in tools), and the client harness
(Claude Code, Codex, IDE plugins, etc.). A fixed benchmark on a fixed codebase with a fixed
agent tells you little about what Serena would add to *your* setup.
- **Predefined tasks bias the measurement.** Choosing specific tasks to evaluate inevitably
introduces selection bias — we would end up picking tasks that either favour or disfavour Serena.
We wanted an evaluation that systematically covers the full surface area of Serena's capabilities
without cherry-picking.
## Design Goals ## Design Goals
Instead of benchmarks, we designed an evaluation methodology with three goals: We designed an evaluation methodology with three goals:
1. **Reproducible by any user, on any project, with any agent.** 1. **Generality.** The evaluation mechanism should be broadly applicable and repeatable by any user, on any project, with any agent.
The evaluation is a single prompt that you give to your agent of choice, pointed at your codebase The evaluation is a single prompt that you give to your agent of choice, pointed at your codebase
of choice, in your client of choice. This means the results directly reflect the value Serena of choice, in your client of choice. This means the results directly reflect the value Serena
would add to your actual workflow — not to an artificial benchmark setup. There is nothing to would add to your actual workflow — not to an artificial benchmark setup. There is nothing to
install, configure, or script beyond what you already have. install, configure, or script beyond what you already have.
2. **The agent evaluates itself.** 2. **The agent evaluates itself.**
We deliberately let the AI agent be both the executor and the evaluator. This may seem We deliberately let the AI agent be both the executor and the evaluator.
counterintuitive, but it is the right design choice: the agent is the actual end user of the Since most evaluations are strictly quantitative, this may seem unusual, but it is a reasonable choice:
The agent is the actual end user of the
tools, so it is in the best position to judge whether a semantic tool improves its workflow tools, so it is in the best position to judge whether a semantic tool improves its workflow
compared to its built-in alternatives. It can measure call counts, payload sizes, and compared to its built-in alternatives. It can measure call counts, payload sizes, and
prerequisite steps from direct experience rather than from proxy metrics. It also avoids the prerequisite steps from direct experience rather than from proxy metrics. It also avoids the
@@ -48,7 +28,7 @@ Instead of benchmarks, we designed an evaluation methodology with three goals:
simply uses them and reports what it observes. simply uses them and reports what it observes.
3. **Comprehensive and unbiased by design.** 3. **Comprehensive and unbiased by design.**
Rather than selecting specific tasks, the prompt defines **task categories** that systematically Rather than selecting specific tasks, the prompt defines *task categories* that systematically
span Serena's capabilities: codebase understanding, single-file edits of varying sizes, multi-file span Serena's capabilities: codebase understanding, single-file edits of varying sizes, multi-file
refactoring, reliability properties, and workflow effects. The agent picks concrete instances refactoring, reliability properties, and workflow effects. The agent picks concrete instances
from the codebase at hand, performs each task using both toolsets side by side, and classifies from the codebase at hand, performs each task using both toolsets side by side, and classifies
@@ -58,7 +38,7 @@ Instead of benchmarks, we designed an evaluation methodology with three goals:
## Method ## Method
We give an AI coding agent a single, detailed [evaluation prompt](010_evaluation-prompt) in a one-shot session. We give an AI coding agent a single, detailed [evaluation prompt](020_prompts/010_evaluation-prompt.md) in a one-shot session.
The prompt instructs the agent to perform approximately 20 hands-on tasks across five areas: The prompt instructs the agent to perform approximately 20 hands-on tasks across five areas:
1. **Codebase understanding** — structural overviews, targeted symbol retrieval, reference finding, type hierarchies, and external dependency lookup. 1. **Codebase understanding** — structural overviews, targeted symbol retrieval, reference finding, type hierarchies, and external dependency lookup.
@@ -72,30 +52,16 @@ tools and its own built-in tools), applies real edits verified via `git diff`, a
payload sizes and prerequisite steps. Edits are reverted after each experiment to keep the working tree clean. payload sizes and prerequisite steps. Edits are reverted after each experiment to keep the working tree clean.
The resulting report classifies each finding into one of three categories: The resulting report classifies each finding into one of three categories:
**(a)** tasks where Serena adds capability,
**(b)** tasks where Serena applies but offers no improvement, and * **(a)** tasks where Serena adds capability,
**(c)** tasks outside Serena's scope. * **(b)** tasks where Serena applies but offers no improvement, and
* **(c)** tasks outside Serena's scope.
Only category (b) constitutes a neutral or negative finding; category (c) is context, not a finding. Only category (b) constitutes a neutral or negative finding; category (c) is context, not a finding.
After the evaluation, a separate [follow-up prompt](011_followup-summary-prompt) asks the agent for a After the evaluation, a separate [follow-up prompt](020_prompts/020_summary-prompt.md) asks the agent for a
one-sentence, user-facing recommendation — the quotes shown on the [main page](https://github.com/oraios/serena). one-sentence, user-facing recommendation — the quotes shown on the [main page](https://github.com/oraios/serena).
## Results
We performed evaluations using popular AI coding agents in representative scenarios — different
agents, different programming languages, and different codebases — to show that the results are
not specific to a single setup.
All evaluations were conducted using the **JetBrains-powered version** of Serena, as it is the
more powerful backend with a broader set of refactoring and navigation capabilities. The
evaluation can easily be repeated with the LSP-based backend to assess its subset of capabilities.
- [Claude Code (Opus 4.6) on a large Python codebase](results/010_cc_on_tianshou-serena-evaluation) — tianshou, a reinforcement learning library (~26K lines).
- [Codex (GPT 5.4) on a Java codebase](results/020_codex_on_jbplugin-serena-evaluation) — the Serena JetBrains plugin itself.
You can run your own evaluation on a project of your choice by reusing our
[evaluation prompt](010_evaluation-prompt) or adapting it to your needs.
## Assessment of the Methodology ## Assessment of the Methodology
_The following assessment was written by Claude Opus 4.6 (high effort) after reading the full _The following assessment was written by Claude Opus 4.6 (high effort) after reading the full
@@ -144,14 +110,12 @@ One thing to watch for is whether agents that are less capable than Opus 4.6 or
meaningful evaluations — the prompt is quite demanding. But that is a question about the agent, not meaningful evaluations — the prompt is quite demanding. But that is a question about the agent, not
about the method. about the method.
## Prompt Fairness ### Prompt Fairness
We put significant effort in crafting an unbiased, fair and objective evaluation prompt. We put significant effort into crafting an unbiased, fair evaluation prompt.
The only "biased" aspects are some sentences about misuse of tools, which we consider irrelevant for the purpose of evaluation To assess the prompt's fairness, we asked Claude Opus 4.6 to evaluate the prompt itself.
(sufficiently intelligent models use the tools correctly).
Here is what Opus 4.6 thinks about our prompt:
**Q**: **Prompt**:
> The attached prompt is designed to evaluate the additive value of Serena as an augmentation layer on top of built-in tools > The attached prompt is designed to evaluate the additive value of Serena as an augmentation layer on top of built-in tools
> (not to compare two competing alternatives). > (not to compare two competing alternatives).
> Is the methodology suitable for that goal, or does it introduce bias that would distort the measurement of what Serena adds? > Is the methodology suitable for that goal, or does it introduce bias that would distort the measurement of what Serena adds?
@@ -160,12 +124,37 @@ Here is what Opus 4.6 thinks about our prompt:
> Do not question the validity of this assumption in your evaluation of the prompt. > Do not question the validity of this assumption in your evaluation of the prompt.
> Give a brief answer, maximally two sentences. > Give a brief answer, maximally two sentences.
> >
> \<Evalution prompt pasted here\> > \<evaluation prompt pasted here\>
**Opus' Answer**: **Claude Opus' Answer**:
> "The methodology is well-suited to its stated goal: by explicitly scoping to correct use, treating Serena as an > "The methodology is well-suited to its stated goal: by explicitly scoping to correct use, treating Serena as an
> augmentation layer rather than a competitor, and classifying out-of-scope tasks as context rather than negative > augmentation layer rather than a competitor, and classifying out-of-scope tasks as context rather than negative
> findings, it avoids the most common biases that would distort a delta measurement. The one risk is that the "correct-use > findings, it avoids the most common biases that would distort a delta measurement. The one risk is that the "correct-use
> rule" could subtly favor Serena by preemptively filtering out any scenario where its tools underperform — but the prompt > rule" could subtly favor Serena by preemptively filtering out any scenario where its tools underperform — but the prompt
> mitigates this by requiring category (b) findings (applies but no improvement) and explicitly demanding negative deltas > mitigates this by requiring category (b) findings (applies but no improvement) and explicitly demanding negative deltas
> be reported, so the bias is acknowledged and structurally counterbalanced." > be reported, so the bias is acknowledged and structurally counterbalanced."
According to Claude, the only "biased" aspects are some sentences about misuse of tools, which we consider irrelevant for the purpose of evaluation
(sufficiently intelligent models use the tools correctly).
## Why Not Benchmarks?
Standard coding benchmarks (SWE-bench, HumanEval, etc.) measure an agent's ability to solve
predefined tasks with a known correct answer. They are valuable for comparing models and agents,
but they are a poor fit for evaluating a **tool augmentation layer** like Serena for several reasons:
- **Benchmarks don't reflect real usage patterns.** Benchmark tasks are typically small, self-contained
problems that can be solved by reading and editing a handful of files. They rarely exercise the
workflows where Serena's tools shine — cross-file refactoring, navigating large codebases by symbol
structure, chaining multiple edits with stable addressing, or querying type hierarchies and
external dependencies. A benchmark score would mostly measure performance on tasks where Serena
is not expected to help.
- **Results would not generalise to the user's project.** Serena's value depends on the codebase
(size, language, complexity), the agent (model, built-in tools), and the client harness
(Claude Code, Codex, IDE plugins, etc.). A fixed benchmark on a fixed codebase with a fixed
agent tells you little about what Serena would add to *your* setup.
- **Predefined tasks bias the measurement.** Choosing specific tasks to evaluate inevitably
introduces selection bias — we would end up picking tasks that either favour or disfavour Serena.
We wanted an evaluation that systematically covers the full surface area of Serena's capabilities
without cherry-picking.
@@ -1,9 +0,0 @@
# Post-Evaluation Summary Prompt
_Use the prompt below to summarize the evaluation of Serena by your agent.
We used it to create the recommendations we display in the readme._
Write a one-sentence user-facing summary about the value that Serena's tools provide for coding. With a bit
of emotion but grounded in the evaluation that you have done above. It should be written from the perspective of a coding AI agent and make it
clear whether an AI agent would ask its owner to add Serena's tools or not. It will be the first sentence that a potential new user will see
about how useful Serena is for an AI agent.
@@ -0,0 +1,6 @@
# Prompts
Our evaluation uses two prompts, which are passed to the LLM in order to generate the evaluation results:
1. the [main evaluation prompt](010_evaluation-prompt), which performs the analysis
2. the [summary prompt](020_summary-prompt), which condenses the results
@@ -1,182 +1,218 @@
# Evaluation Prompt # Evaluation Prompt
Use the prompt below to evaluate the added value of Serena's tools against your agent's built-in tools on a project of We use the prompt below to evaluate the added value of Serena's tools against the
your choice. agent's built-in tools on a given project.
All evaluations that you find in our documentation were created in one-shot sessions, only using this prompt and The evaluations were created in one-shot sessions, only using this
then following up with a separate [summary prompt](011_followup-summary-prompt) prompt and the follow-up [prompt for summarization](020_summary-prompt)
```
# Evaluate Serena's Tools Against Built-Ins # Evaluate Serena's Tools Against Built-Ins
You have access to Serena's coding tools alongside your built-in tools (Read, Edit, Write, Glob, Grep, Bash, etc.). I You have access to Serena's coding tools alongside your built-in tools (Read,
want a thorough, evidence-based evaluation of **what Serena's tools add on top of the built-ins**, assuming both Edit, Write, Glob, Grep, Bash, etc.). I want a thorough, evidence-based
evaluation of **what Serena's tools add on top of the built-ins**, assuming both
toolsets are used correctly. toolsets are used correctly.
This is an evaluation, not a user guide, and it is not a binary adoption pitch. Your job is to answer: *if a competent This is an evaluation, not a user guide, and it is not a binary adoption pitch.
user of both toolsets had only the built-ins, what concrete capabilities and efficiency differences would they Your job is to answer: *if a competent user of both toolsets had only the
experience, and by how much?* A reader who finishes your report should have a clear, specific picture of what Serena built-ins, what concrete capabilities and efficiency differences would they
adds — as well as where it provides no meaningful improvement or introduces tradeoffs — in terms of capabilities, experience, and by how much?* A reader who finishes your report should have a
workflows, and efficiency. Not a thumbs-up/thumbs-down, but a sharp description of the delta. clear, specific picture of what Serena adds — as well as where it provides no
meaningful improvement or introduces tradeoffs — in terms of capabilities,
workflows, and efficiency. Not a thumbs-up/thumbs-down, but a sharp description
of the delta.
Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be careful of X" warnings are out of scope. Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be
They belong in onboarding material for a developer learning the tools, not in a delta analysis of what the tools add. careful of X" warnings are out of scope. They belong in onboarding material for
a developer learning the tools, not in a delta analysis of what the tools add.
**Describe the measured differences clearly and neutrally.** Avoid generic or non-informative framing such as "both have **Describe the measured differences clearly and neutrally.** Avoid generic or
their place" unless supported by concrete findings. If Serena adds substantial capabilities, name and quantify them. If non-informative framing such as "both have their place" unless supported by
it adds marginal or no capabilities, say that and show why. If there are regressions or tradeoffs, include them concrete findings. If Serena adds substantial capabilities, name and quantify
explicitly. The two toolsets are complementary — that's a given, not the answer. them. If it adds marginal or no capabilities, say that and show why. If there
Serena is an augmentation layer, not a replacement. Do not penalize it for tasks it was not designed to address — are regressions or tradeoffs, include them explicitly. The two toolsets are
instead, note those tasks as "built-in only" and move on. The evaluation should measure what Serena adds where it complementary — that's a given, not the answer. Serena is an augmentation layer,
applies, not what it fails to add where it doesn't.The answer is a specific list of what not a replacement. Do not penalize it for tasks it was not designed to address —
Serena contributes (or does not contribute) to a correct-use workflow relative to built-ins. instead, note those tasks as "built-in only" and move on. The evaluation should
measure what Serena adds where it applies, not what it fails to add where it
doesn't.The answer is a specific list of what Serena contributes (or does not
contribute) to a correct-use workflow relative to built-ins.
Write the report to serena-evaluation.md in the repo root. Write the report to serena-evaluation.md in the repo root.
---
## Ground rules ## Ground rules
### Starting conditions ### Starting conditions
- Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes about the repo. Do not read - Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes
documentation files either. Explore as if you've never seen it, focusing on code. about the repo. Do not read documentation files either. Explore as if you've
- Use git as your safety net — experiment freely. Any edit can be reverted with `git checkout -- <file>` or `git stash`. never seen it, focusing on code.
Run edits for real; don't simulate. A hands-on comparison is worth far more than a thought experiment. - Use git as your safety net — experiment freely. Any edit can be reverted with
- After each experiment, verify the working tree is clean with `git status --short` before moving on. `git checkout -- <file>` or `git stash`. Run edits for real; don't simulate. A
hands-on comparison is worth far more than a thought experiment.
- After each experiment, verify the working tree is clean with
`git status --short` before moving on.
### How to compare — correct use only ### How to compare — correct use only
- **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed for, called the way a competent user - **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed
would call it. A tool doing exactly what its contract says is not a finding, even if a careless caller could misuse for, called the way a competent user would call it. A tool doing exactly what
it. its contract says is not a finding, even if a careless caller could misuse it.
- **Know the contract before you call.** Before invoking any tool, have a one-sentence understanding of what it does. If - **Know the contract before you call.** Before invoking any tool, have a
you expect an error or "not applicable," don't make the call. one-sentence understanding of what it does. If you expect an error or "not
- **Refactoring semantics are real.** Inlining requires a substitutable function (typically single-expression, no side applicable," don't make the call.
effects); moving requires a legal target; safe-delete requires no surviving usages. If the repo has no suitable - **Refactoring semantics are real.** Inlining requires a substitutable function
candidate for a given refactoring, report "no suitable candidate in this codebase" and skip it — don't contrive a (typically single-expression, no side effects); moving requires a legal
broken input. target; safe-delete requires no surviving usages. If the repo has no suitable
candidate for a given refactoring, report "no suitable candidate in this
codebase" and skip it — don't contrive a broken input.
### How to compare — workflow level, not single-call level ### How to compare — workflow level, not single-call level
- For every task, write out the full end-to-end call chain on each side before drawing conclusions. Include prerequisite - For every task, write out the full end-to-end call chain on each side before
reads and follow-up steps. drawing conclusions. Include prerequisite reads and follow-up steps.
- Do not evaluate a tool based on criteria that only arise from mixing workflows incorrectly. - Do not evaluate a tool based on criteria that only arise from mixing workflows
- Ephemeral addressing is a liability. Line numbers and byte offsets go stale after edits; stable addressing (name incorrectly.
paths) may reduce rework. - Ephemeral addressing is a liability. Line numbers and byte offsets go stale
after edits; stable addressing (name paths) may reduce rework.
### How to measure ### How to measure
- Track observations during execution. For every tool call, note: number of calls, approximate input size, output size, - Track observations during execution. For every tool call, note: number of
and any prerequisite or verification steps. calls, approximate input size, output size, and any prerequisite or
- Separate call count, input payload, output payload, and verification cost as distinct axes. verification steps.
- Separate call count, input payload, output payload, and verification cost as
distinct axes.
- Include prerequisite Reads and post-hoc verification steps in comparisons. - Include prerequisite Reads and post-hoc verification steps in comparisons.
When a task falls entirely outside Serena's design scope (e.g., reading config files, small text edits where Edit When a task falls entirely outside Serena's design scope (e.g., reading config
already sends minimal payload), classify it as "not applicable" rather than as a negative delta. A negative delta files, small text edits where Edit already sends minimal payload), classify it
requires that Serena targets the task and performs worse, not that a tool designed for something else is suboptimal when as "not applicable" rather than as a negative delta. A negative delta requires
misapplied to it. that Serena targets the task and performs worse, not that a tool designed for
something else is suboptimal when misapplied to it.
---
## Exploration phase — tasks to actually perform ## Exploration phase — tasks to actually perform
Work through the following. Each item exercises a specific capability under correct use; substitute an equivalent if an Work through the following. Each item exercises a specific capability under
item isn't applicable. correct use; substitute an equivalent if an item isn't applicable.
### Codebase understanding ### Codebase understanding
1. Get a high-level overview of the repository structure — top-level layout, main packages, entry points. 1. Get a high-level overview of the repository structure — top-level layout,
2. Pick one large source file (300+ lines). Get a structural overview of it. Do it with semantic overview tools and with main packages, entry points.
Glob/Grep/Read. Then write out the concrete next step on each side and compare the pair of calls, not just the 2. Pick one large source file (300+ lines). Get a structural overview of it. Do
it with semantic overview tools and with Glob/Grep/Read. Then write out the
concrete next step on each side and compare the pair of calls, not just the
overview call. overview call.
3. Pick a specific method inside a class and retrieve its body without reading the surrounding file. 3. Pick a specific method inside a class and retrieve its body without reading
4. For one non-trivial symbol, find all references across the codebase. Compare recall and precision under the the surrounding file.
question "who uses this in code?" vs "where is this mentioned anywhere, including docs?" 4. For one non-trivial symbol, find all references across the codebase. Compare
5. For a class, list its subclasses / implementations and its supertypes, including transitively. Compare against what recall and precision under the question "who uses this in code?" vs "where is
text search would need to do. this mentioned anywhere, including docs?"
6. For at least one symbol from an external dependency (a third-party library), try to retrieve its definition or 5. For a class, list its subclasses / implementations and its supertypes,
signature. Note whether each toolset can do this at all and what infrastructure it requires (environment activation, including transitively. Compare against what text search would need to do.
6. For at least one symbol from an external dependency (a third-party library),
try to retrieve its definition or signature. Note whether each toolset can do
this at all and what infrastructure it requires (environment activation,
site-packages discovery, language-server indexing, etc.). site-packages discovery, language-server indexing, etc.).
### Single-file edits — span the full range of edit sizes ### Single-file edits — span the full range of edit sizes
7a. Small tweak (1–3 lines inside a method). Change an error message or rename a local variable inside a larger method. 7a. Small tweak (1–3 lines inside a method). Change an error message or rename a
Do it with `Edit` and with symbolic body replacement. Compare payload sent, payload received, and prerequisite reads. local variable inside a larger method. Do it with `Edit` and with symbolic body
replacement. Compare payload sent, payload received, and prerequisite reads.
7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the main logic of a method while keeping its 7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the
signature. Do it both ways. main logic of a method while keeping its signature. Do it both ways.
7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire body. Do it both ways. 7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire
body. Do it both ways.
8. Insert a new function/method at a specific structural location (for example, right after an existing method). Try 8. Insert a new function/method at a specific structural location (for example,
both the symbolic-insert path and the manual Edit path. right after an existing method). Try both the symbolic-insert path and the
9. Rename a private helper used only within one file. Compare doing it by hand vs. using a semantic rename. manual Edit path.
9. Rename a private helper used only within one file. Compare doing it by hand
vs. using a semantic rename.
### Multi-file changes ### Multi-file changes
10. Rename a symbol (function, class, or method) used across several files including imports. Compare the semantic path 10. Rename a symbol (function, class, or method) used across several files
against the built-in equivalent chain. including imports. Compare the semantic path against the built-in equivalent
11. Move a symbol from one module to another, updating imports at all call sites. Use the semantic move tool if chain.
available; plan the built-in equivalent honestly. 11. Move a symbol from one module to another, updating imports at all call
12. Move a file or package to a different location, updating imports at all call sites. Use the semantic move tool if sites. Use the semantic move tool if available; plan the built-in equivalent
available; plan the built-in equivalent honestly. honestly.
12. Delete a symbol safely, checking it has no remaining usages. Compare search-then-delete with a safe-delete tool. 12. Move a file or package to a different location, updating imports at all call
13. Delete a symbol and propagate the deletion to all call sites. Compare to how the built-in equivalent would work. sites. Use the semantic move tool if available; plan the built-in equivalent
13. Inline a small helper into its call sites — only if the codebase contains a function that is legally inlinable. If honestly.
no such candidate exists, report "no suitable candidate" and skip it. 12. Delete a symbol safely, checking it has no remaining usages. Compare
search-then-delete with a safe-delete tool.
13. Delete a symbol and propagate the deletion to all call sites. Compare to how
the built-in equivalent would work.
13. Inline a small helper into its call sites — only if the codebase contains a
function that is legally inlinable. If no such candidate exists, report "no
suitable candidate" and skip it.
### Reliability & correctness under correct use ### Reliability & correctness under correct use
14. Scope precision. Demonstrate that semantic tools address symbols by name path and can target a specific class 14. Scope precision. Demonstrate that semantic tools address symbols by name
method, override, or overload that text search would over-match. path and can target a specific class method, override, or overload that text
15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are updated or none. A chain of `Edit` search would over-match.
calls is not. 15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are
16. Success signals. For each completed refactor, note what each tool returns on success. updated or none. A chain of `Edit` calls is not.
16. Success signals. For each completed refactor, note what each tool returns on
success.
### Workflow effects across multiple edits ### Workflow effects across multiple edits
17. Chain at least three edits in one file. Report what each toolset requires between edits. 17. Chain at least three edits in one file. Report what each toolset requires
18. Multi-step exploration across the repo. Note whether intermediate results remain useful across later edits or have between edits.
to be refreshed. 18. Multi-step exploration across the repo. Note whether intermediate results
remain useful across later edits or have to be refreshed.
### Things where the comparison shouldn't be interesting ### Things where the comparison shouldn't be interesting
19. Read and understand a non-code file (config, changelog, docs, notebook). Semantic-code tools don't apply — use 19. Read and understand a non-code file (config, changelog, docs, notebook).
`Read`. Semantic-code tools don't apply — use `Read`.
20. Search for a free-text pattern across the repo (log string, magic constant, URL). Use `Grep`. 20. Search for a free-text pattern across the repo (log string, magic constant,
URL). Use `Grep`.
---
## Evaluation phase ## Evaluation phase
Write a report structured for progressive disclosure. Write a report structured for progressive disclosure.
**Value-weighting is required.** For every contribution or difference you identify — positive, neutral, or negative — **Value-weighting is required.** For every contribution or difference you
estimate: identify — positive, neutral, or negative — estimate:
- **Frequency:** how often this arises in typical coding work - **Frequency:** how often this arises in typical coding work
- **Value per hit:** calls saved, tokens saved, or correctness impact - **Value per hit:** calls saved, tokens saved, or correctness impact
Order findings by **frequency × value-per-hit**, not novelty. Order findings by **frequency × value-per-hit**, not novelty.
**Every section must end with a one-sentence verdict** summarizing the practical takeaway. **Every section must end with a one-sentence verdict** summarizing the practical
takeaway.
---
### 1. Headline: what Serena changes ### 1. Headline: what Serena changes
Open with a precise description of the delta Serena provides. Open with a precise description of the delta Serena provides. Distinguish
Distinguish between three categories: between three categories: (a) tasks where Serena adds capability, (b) tasks
(a) tasks where Serena adds capability, where Serena applies but offers no improvement, and (c) tasks outside Serena's
(b) tasks where Serena applies but offers scope. Only category (b) constitutes a neutral or negative finding. Category (c)
no improvement, and (c) tasks outside Serena's scope. is context, not a finding.
Only category (b) constitutes a neutral or negative finding.
Category (c) is context, not a finding.
A reader stopping here should understand both what is gained and what is not. A reader stopping here should understand both what is gained and what is not.
**Verdict:** (one sentence) **Verdict:** (one sentence)
---
### 2. Added value and differences by area (3–6 bullets) ### 2. Added value and differences by area (3–6 bullets)
@@ -190,7 +226,7 @@ Avoid framing in terms of “wins”; describe concrete differences.
**Verdict:** (one sentence) **Verdict:** (one sentence)
---
### 3. Detailed evidence, grouped by capability ### 3. Detailed evidence, grouped by capability
@@ -208,7 +244,7 @@ Include cases where:
End each subsection with a verdict. End each subsection with a verdict.
---
### 4. Token-efficiency analysis ### 4. Token-efficiency analysis
@@ -222,7 +258,7 @@ Include cases where each toolset is more efficient.
**Verdict:** (one sentence) **Verdict:** (one sentence)
---
### 5. Reliability & correctness (under correct use) ### 5. Reliability & correctness (under correct use)
@@ -238,35 +274,36 @@ Include both strengths and limitations of each toolset.
**Verdict:** (one sentence) **Verdict:** (one sentence)
---
### 6. Workflow effects across a session ### 6. Workflow effects across a session
Evaluate multi-step workflows and whether advantages compound or diminish. Include neutral or negative findings where Evaluate multi-step workflows and whether advantages compound or diminish.
applicable. Include neutral or negative findings where applicable.
**Verdict:** (one sentence) **Verdict:** (one sentence)
---
### 7. Unique capabilities (if any) ### 7. Unique capabilities (if any)
List capabilities that have no practical built-in equivalent. If none exist, explicitly state that. Annotate each with List capabilities that have no practical built-in equivalent. If none exist,
frequency and impact. explicitly state that. Annotate each with frequency and impact.
**Verdict:** (one sentence) **Verdict:** (one sentence)
---
### 8. Tasks outside Serena's scope (built-in only) ### 8. Tasks outside Serena's scope (built-in only)
Identify tasks where built-ins are the natural choice because Serena's tools don't target them. List these briefly for Identify tasks where built-ins are the natural choice because Serena's tools
completeness but do not frame them as Serena shortcomings — they are outside its scope. Estimate their share of daily don't target them. List these briefly for completeness but do not frame them as
Serena shortcomings — they are outside its scope. Estimate their share of daily
work to contextualize how much of a session Serena's augmentation covers. work to contextualize how much of a session Serena's augmentation covers.
**Verdict:** (one sentence) **Verdict:** (one sentence)
---
### 9. Practical usage rule ### 9. Practical usage rule
@@ -274,7 +311,7 @@ Provide a decision rule for choosing between toolsets based on task type.
**Verdict:** (one sentence) **Verdict:** (one sentence)
---
## What I'm looking for ## What I'm looking for
@@ -283,7 +320,7 @@ Provide a decision rule for choosing between toolsets based on task type.
- Clear quantification of impact - Clear quantification of impact
- Honest workflow-level comparisons - Honest workflow-level comparisons
---
## What I am not looking for ## What I am not looking for
@@ -292,4 +329,5 @@ Provide a decision rule for choosing between toolsets based on task type.
- Neutral statements without evidence - Neutral statements without evidence
- Binary recommendations - Binary recommendations
- Novelty-driven ordering - Novelty-driven ordering
- Unquantified claims - Unquantified claims
```
@@ -0,0 +1,12 @@
# Summary Prompt
We used the prompt below to summarize the evaluation of Serena.
```
Write a one-sentence user-facing summary about the value that Serena's tools
provide for coding. With a bit of emotion but grounded in the evaluation that
you have done above. It should be written from the perspective of a coding AI
agent and make it clear whether an AI agent would ask its owner to add Serena's
tools or not. It will be the first sentence that a potential new user will see
about how useful Serena is for an AI agent.
```
@@ -0,0 +1,17 @@
# Results
This section presents the results of the evaluation.
We performed evaluations using popular AI coding agents in representative scenarios — different
agents, different programming languages, and different codebases — to show that the results are
not specific to a single setup.
All evaluations were conducted using the **JetBrains-powered version** of Serena, as it is the
more powerful backend with a broader set of refactoring and navigation capabilities. The
evaluation can easily be repeated with the LSP-based backend to assess its subset of capabilities.
- [Claude Code (Opus 4.6) on a large Python codebase](010_cc_on_tianshou)
- [Codex (GPT 5.4) on a Java codebase](020_codex_on_jbplugin)
You can run your own evaluation on a project of your choice by reusing our
[evaluation prompt](../020_prompts/010_evaluation-prompt.md).
@@ -1,9 +1,13 @@
# Evaluation Report: Serena's Tools vs Built-In Tools :::{admonition} Evaluation Result
:class: note
**Generated by**: Claude Opus 4.6 (coding AI agent in Claude Code CLI)
**Codebase:** [Tianshou](https://github.com/thu-ml/tianshou) — a Python reinforcement learning library (~26K lines, 43 source files)
:::
# Claude Code (Opus)
> **One-line summary:** Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up. > **One-line summary:** Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up.
**Evaluated by:** Claude Opus 4.6 (coding AI agent in Claude Code CLI)
**Codebase:** [tianshou](https://github.com/thu-ml/tianshou) — a Python reinforcement learning library (~26K lines, 43 source files)
**Method:** Hands-on, side-by-side execution of 20 tasks using both toolsets. All edits were applied to real files and verified via `git diff`, then reverted. **Method:** Hands-on, side-by-side execution of 20 tasks using both toolsets. All edits were applied to real files and verified via `git diff`, then reverted.
--- ---
@@ -1,14 +1,19 @@
# Evaluation: What Serena Adds Over Built-Ins :::{admonition} Evaluation Result
:class: note
**Generated by:** GPT-5.4 (high) in Codex
**Codebase:** Serena JetBrains Plugin (Java)
:::
**Evaluated by:** Gpt 5.4 (high) in Codex # Codex (GPT-5.4)
> **One-line summary:** As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and
refactoring, which makes real code changes feel faster, safer, and far less blind.
This report compares Serena's JetBrains-backed semantic coding tools with built-in file, shell, search, and patch tools in this repository. The comparison assumes competent use of both toolsets: built-ins are used for text, file, shell, config, and small patch work; Serena is used where symbol identity, language semantics, or IDE refactoring semantics apply. This report compares Serena's JetBrains-backed semantic coding tools with built-in file, shell, search, and patch tools in this repository. The comparison assumes competent use of both toolsets: built-ins are used for text, file, shell, config, and small patch work; Serena is used where symbol identity, language semantics, or IDE refactoring semantics apply.
Method: I explored code first, avoided repo documentation and prior notes, ran real edits/refactors, and reverted after each experiment. After each edit/refactor experiment I checked `git status --short` and returned the tree to clean before moving on. Measurements are approximate, but call counts, diff sizes, and result shapes are from observed runs. Method: I explored code first, avoided repo documentation and prior notes, ran real edits/refactors, and reverted after each experiment. After each edit/refactor experiment I checked `git status --short` and returned the tree to clean before moving on. Measurements are approximate, but call counts, diff sizes, and result shapes are from observed runs.
> **One-line summary:** As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and
refactoring, which makes real code changes feel faster, safer, and far less blind.
## 1. Headline: what Serena changes ## 1. Headline: what Serena changes
Serena adds a semantic layer over the codebase. Its concrete delta is the ability to address and transform code by symbols: name paths, overload indexes, reference graphs, type hierarchies, external declarations, and JetBrains refactoring operations. Serena adds a semantic layer over the codebase. Its concrete delta is the ability to address and transform code by symbols: name paths, overload indexes, reference graphs, type hierarchies, external declarations, and JetBrains refactoring operations.
-3
View File
@@ -1,3 +0,0 @@
# Results
Here are the results of the evaluation for various coding agents running on various projects.
+1 -1
View File
@@ -102,7 +102,7 @@ sphinx:
local_extensions : # A list of local extensions to load by sphinx specified by "name: path" items local_extensions : # A list of local extensions to load by sphinx specified by "name: path" items
recursive_update : false # A boolean indicating whether to overwrite the Sphinx config (true) or recursively update (false) recursive_update : false # A boolean indicating whether to overwrite the Sphinx config (true) or recursively update (false)
config : # key-value pairs to directly over-ride the Sphinx configuration config : # key-value pairs to directly over-ride the Sphinx configuration
master_doc: "01-about/000_intro.md" master_doc: "01-about/000_evaluation-results.md"
html_theme_options: html_theme_options:
logo: logo:
image_light: ../resources/serena-logo.svg image_light: ../resources/serena-logo.svg
+1 -1
View File
@@ -190,7 +190,7 @@ def autogen_about_intro_features():
autogen_info = f"<!-- This section is auto-generated by {__file__} from the root README.md; do not edit. -->\n\n" autogen_info = f"<!-- This section is auto-generated by {__file__} from the root README.md; do not edit. -->\n\n"
with open(Path(__file__).parent / "01-about" / "000_intro.md", "w", encoding="utf-8") as f: with open(Path(__file__).parent / "01-about" / "000_evaluation-results.md", "w", encoding="utf-8") as f:
f.write(autogen_info) f.write(autogen_info)
f.write("# About Serena\n\n") f.write("# About Serena\n\n")
f.write(f"**{tagline}**\n\n") f.write(f"**{tagline}**\n\n")