Restructure evaluation docs

This commit is contained in:
Dominik Jain committed 2026-04-13 16:42:57 +02:00
1 parent c3a5652fc7
commit 46be9c4bd7
12 files changed
+260 -199

No files matched your search

@@ -0,0 +1,2 @@
# Evaluation
@@ -1,46 +1,26 @@
# Evaluation
# Methodology
In this section we describe how we evaluate the performance of Serena's tools.
In this section we describe the methodology we applied in evaluating the performance of Serena's tools.
The evaluation measures the **concrete delta** that Serena's tools provide on top of an agent's
built-in capabilities (file reads, text edits, grep, shell, etc.).
Rather than a simple thumbs-up/thumbs-down, it produces a detailed, evidence-based report
covering capabilities, efficiency and reliability.
## Why Not Benchmarks?
Standard coding benchmarks (SWE-bench, HumanEval, etc.) measure an agent's ability to solve
predefined tasks with a known correct answer. They are valuable for comparing models and agents,
but they are a poor fit for evaluating a **tool augmentation layer** like Serena for several reasons:
- **Benchmarks don't reflect real usage patterns.** Benchmark tasks are typically small, self-contained
problems that can be solved by reading and editing a handful of files. They rarely exercise the
workflows where Serena's tools shine — cross-file refactoring, navigating large codebases by symbol
structure, chaining multiple edits with stable addressing, or querying type hierarchies and
external dependencies. A benchmark score would mostly measure performance on tasks where Serena
is not expected to help.
- **Results would not generalise to the user's project.** Serena's value depends on the codebase
(size, language, complexity), the agent (model, built-in tools), and the client harness
(Claude Code, Codex, IDE plugins, etc.). A fixed benchmark on a fixed codebase with a fixed
agent tells you little about what Serena would add to *your* setup.
- **Predefined tasks bias the measurement.** Choosing specific tasks to evaluate inevitably
introduces selection bias — we would end up picking tasks that either favour or disfavour Serena.
We wanted an evaluation that systematically covers the full surface area of Serena's capabilities
without cherry-picking.
## Design Goals
Instead of benchmarks, we designed an evaluation methodology with three goals:
We designed an evaluation methodology with three goals:
1. **Reproducible by any user, on any project, with any agent.**
1. **Generality.** The evaluation mechanism should be broadly applicable and repeatable by any user, on any project, with any agent.
The evaluation is a single prompt that you give to your agent of choice, pointed at your codebase
of choice, in your client of choice. This means the results directly reflect the value Serena
would add to your actual workflow — not to an artificial benchmark setup. There is nothing to
install, configure, or script beyond what you already have.
2. **The agent evaluates itself.**
We deliberately let the AI agent be both the executor and the evaluator. This may seem
counterintuitive, but it is the right design choice: the agent is the actual end user of the
We deliberately let the AI agent be both the executor and the evaluator.
Since most evaluations are strictly quantitative, this may seem unusual, but it is a reasonable choice:
The agent is the actual end user of the
tools, so it is in the best position to judge whether a semantic tool improves its workflow
compared to its built-in alternatives. It can measure call counts, payload sizes, and
prerequisite steps from direct experience rather than from proxy metrics. It also avoids the
@@ -48,7 +28,7 @@ Instead of benchmarks, we designed an evaluation methodology with three goals:
simply uses them and reports what it observes.
3. **Comprehensive and unbiased by design.**
Rather than selecting specific tasks, the prompt defines **task categories** that systematically
Rather than selecting specific tasks, the prompt defines *task categories* that systematically
span Serena's capabilities: codebase understanding, single-file edits of varying sizes, multi-file
refactoring, reliability properties, and workflow effects. The agent picks concrete instances
from the codebase at hand, performs each task using both toolsets side by side, and classifies
@@ -58,7 +38,7 @@ Instead of benchmarks, we designed an evaluation methodology with three goals:
## Method
We give an AI coding agent a single, detailed [evaluation prompt](010_evaluation-prompt) in a one-shot session.
We give an AI coding agent a single, detailed [evaluation prompt](020_prompts/010_evaluation-prompt.md) in a one-shot session.
The prompt instructs the agent to perform approximately 20 hands-on tasks across five areas:
1. **Codebase understanding** — structural overviews, targeted symbol retrieval, reference finding, type hierarchies, and external dependency lookup.
@@ -72,30 +52,16 @@ tools and its own built-in tools), applies real edits verified via `git diff`, a
payload sizes and prerequisite steps. Edits are reverted after each experiment to keep the working tree clean.
The resulting report classifies each finding into one of three categories:
**(a)** tasks where Serena adds capability,
**(b)** tasks where Serena applies but offers no improvement, and
**(c)** tasks outside Serena's scope.
* **(a)** tasks where Serena adds capability,
* **(b)** tasks where Serena applies but offers no improvement, and
* **(c)** tasks outside Serena's scope.
Only category (b) constitutes a neutral or negative finding; category (c) is context, not a finding.
After the evaluation, a separate [follow-up prompt](011_followup-summary-prompt) asks the agent for a
After the evaluation, a separate [follow-up prompt](020_prompts/020_summary-prompt.md) asks the agent for a
one-sentence, user-facing recommendation — the quotes shown on the [main page](https://github.com/oraios/serena).
## Results
We performed evaluations using popular AI coding agents in representative scenarios — different
agents, different programming languages, and different codebases — to show that the results are
not specific to a single setup.
All evaluations were conducted using the **JetBrains-powered version** of Serena, as it is the
more powerful backend with a broader set of refactoring and navigation capabilities. The
evaluation can easily be repeated with the LSP-based backend to assess its subset of capabilities.
- [Claude Code (Opus 4.6) on a large Python codebase](results/010_cc_on_tianshou-serena-evaluation) — tianshou, a reinforcement learning library (~26K lines).
- [Codex (GPT 5.4) on a Java codebase](results/020_codex_on_jbplugin-serena-evaluation) — the Serena JetBrains plugin itself.
You can run your own evaluation on a project of your choice by reusing our
[evaluation prompt](010_evaluation-prompt) or adapting it to your needs.
## Assessment of the Methodology
_The following assessment was written by Claude Opus 4.6 (high effort) after reading the full
@@ -144,14 +110,12 @@ One thing to watch for is whether agents that are less capable than Opus 4.6 or
meaningful evaluations — the prompt is quite demanding. But that is a question about the agent, not
about the method.
## Prompt Fairness
### Prompt Fairness
We put significant effort in crafting an unbiased, fair and objective evaluation prompt.
The only "biased" aspects are some sentences about misuse of tools, which we consider irrelevant for the purpose of evaluation
(sufficiently intelligent models use the tools correctly).
Here is what Opus 4.6 thinks about our prompt:
We put significant effort into crafting an unbiased, fair evaluation prompt.
To assess the prompt's fairness, we asked Claude Opus 4.6 to evaluate the prompt itself.
**Q**:
**Prompt**:
> The attached prompt is designed to evaluate the additive value of Serena as an augmentation layer on top of built-in tools
> (not to compare two competing alternatives).
> Is the methodology suitable for that goal, or does it introduce bias that would distort the measurement of what Serena adds?
@@ -160,12 +124,37 @@ Here is what Opus 4.6 thinks about our prompt:
> Do not question the validity of this assumption in your evaluation of the prompt.
> Give a brief answer, maximally two sentences.
>
> \<Evalution prompt pasted here\>
> \<evaluation prompt pasted here\>
**Opus' Answer**:
**Claude Opus' Answer**:
> "The methodology is well-suited to its stated goal: by explicitly scoping to correct use, treating Serena as an
> augmentation layer rather than a competitor, and classifying out-of-scope tasks as context rather than negative
> findings, it avoids the most common biases that would distort a delta measurement. The one risk is that the "correct-use
> rule" could subtly favor Serena by preemptively filtering out any scenario where its tools underperform — but the prompt
> mitigates this by requiring category (b) findings (applies but no improvement) and explicitly demanding negative deltas
> be reported, so the bias is acknowledged and structurally counterbalanced."
> be reported, so the bias is acknowledged and structurally counterbalanced."
According to Claude, the only "biased" aspects are some sentences about misuse of tools, which we consider irrelevant for the purpose of evaluation
(sufficiently intelligent models use the tools correctly).
## Why Not Benchmarks?
Standard coding benchmarks (SWE-bench, HumanEval, etc.) measure an agent's ability to solve
predefined tasks with a known correct answer. They are valuable for comparing models and agents,
but they are a poor fit for evaluating a **tool augmentation layer** like Serena for several reasons:
- **Benchmarks don't reflect real usage patterns.** Benchmark tasks are typically small, self-contained
problems that can be solved by reading and editing a handful of files. They rarely exercise the
workflows where Serena's tools shine — cross-file refactoring, navigating large codebases by symbol
structure, chaining multiple edits with stable addressing, or querying type hierarchies and
external dependencies. A benchmark score would mostly measure performance on tasks where Serena
is not expected to help.
- **Results would not generalise to the user's project.** Serena's value depends on the codebase
(size, language, complexity), the agent (model, built-in tools), and the client harness
(Claude Code, Codex, IDE plugins, etc.). A fixed benchmark on a fixed codebase with a fixed
agent tells you little about what Serena would add to *your* setup.
- **Predefined tasks bias the measurement.** Choosing specific tasks to evaluate inevitably
introduces selection bias — we would end up picking tasks that either favour or disfavour Serena.
We wanted an evaluation that systematically covers the full surface area of Serena's capabilities
without cherry-picking.
@@ -1,9 +0,0 @@
# Post-Evaluation Summary Prompt
_Use the prompt below to summarize the evaluation of Serena by your agent.
We used it to create the recommendations we display in the readme._
Write a one-sentence user-facing summary about the value that Serena's tools provide for coding. With a bit
of emotion but grounded in the evaluation that you have done above. It should be written from the perspective of a coding AI agent and make it
clear whether an AI agent would ask its owner to add Serena's tools or not. It will be the first sentence that a potential new user will see
about how useful Serena is for an AI agent.
@@ -0,0 +1,6 @@
# Prompts
Our evaluation uses two prompts, which are passed to the LLM in order to generate the evaluation results:
1. the [main evaluation prompt](010_evaluation-prompt), which performs the analysis
2. the [summary prompt](020_summary-prompt), which condenses the results
@@ -1,182 +1,218 @@
# Evaluation Prompt
Use the prompt below to evaluate the added value of Serena's tools against your agent's built-in tools on a project of
your choice.
All evaluations that you find in our documentation were created in one-shot sessions, only using this prompt and
then following up with a separate [summary prompt](011_followup-summary-prompt)
We use the prompt below to evaluate the added value of Serena's tools against the
agent's built-in tools on a given project.
The evaluations were created in one-shot sessions, only using this
prompt and the follow-up [prompt for summarization](020_summary-prompt)
```
# Evaluate Serena's Tools Against Built-Ins
You have access to Serena's coding tools alongside your built-in tools (Read, Edit, Write, Glob, Grep, Bash, etc.). I
want a thorough, evidence-based evaluation of **what Serena's tools add on top of the built-ins**, assuming both
You have access to Serena's coding tools alongside your built-in tools (Read,
Edit, Write, Glob, Grep, Bash, etc.). I want a thorough, evidence-based
evaluation of **what Serena's tools add on top of the built-ins**, assuming both
toolsets are used correctly.
This is an evaluation, not a user guide, and it is not a binary adoption pitch. Your job is to answer: *if a competent
user of both toolsets had only the built-ins, what concrete capabilities and efficiency differences would they
experience, and by how much?* A reader who finishes your report should have a clear, specific picture of what Serena
adds — as well as where it provides no meaningful improvement or introduces tradeoffs — in terms of capabilities,
workflows, and efficiency. Not a thumbs-up/thumbs-down, but a sharp description of the delta.
This is an evaluation, not a user guide, and it is not a binary adoption pitch.
Your job is to answer: *if a competent user of both toolsets had only the
built-ins, what concrete capabilities and efficiency differences would they
experience, and by how much?* A reader who finishes your report should have a
clear, specific picture of what Serena adds — as well as where it provides no
meaningful improvement or introduces tradeoffs — in terms of capabilities,
workflows, and efficiency. Not a thumbs-up/thumbs-down, but a sharp description
of the delta.
Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be careful of X" warnings are out of scope.
They belong in onboarding material for a developer learning the tools, not in a delta analysis of what the tools add.
Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be
careful of X" warnings are out of scope. They belong in onboarding material for
a developer learning the tools, not in a delta analysis of what the tools add.
**Describe the measured differences clearly and neutrally.** Avoid generic or non-informative framing such as "both have
their place" unless supported by concrete findings. If Serena adds substantial capabilities, name and quantify them. If
it adds marginal or no capabilities, say that and show why. If there are regressions or tradeoffs, include them
explicitly. The two toolsets are complementary — that's a given, not the answer.
Serena is an augmentation layer, not a replacement. Do not penalize it for tasks it was not designed to address —
instead, note those tasks as "built-in only" and move on. The evaluation should measure what Serena adds where it
applies, not what it fails to add where it doesn't.The answer is a specific list of what
Serena contributes (or does not contribute) to a correct-use workflow relative to built-ins.
**Describe the measured differences clearly and neutrally.** Avoid generic or
non-informative framing such as "both have their place" unless supported by
concrete findings. If Serena adds substantial capabilities, name and quantify
them. If it adds marginal or no capabilities, say that and show why. If there
are regressions or tradeoffs, include them explicitly. The two toolsets are
complementary — that's a given, not the answer. Serena is an augmentation layer,
not a replacement. Do not penalize it for tasks it was not designed to address —
instead, note those tasks as "built-in only" and move on. The evaluation should
measure what Serena adds where it applies, not what it fails to add where it
doesn't.The answer is a specific list of what Serena contributes (or does not
contribute) to a correct-use workflow relative to built-ins.
Write the report to serena-evaluation.md in the repo root.
---
## Ground rules
### Starting conditions
- Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes about the repo. Do not read
documentation files either. Explore as if you've never seen it, focusing on code.
- Use git as your safety net — experiment freely. Any edit can be reverted with `git checkout -- <file>` or `git stash`.
Run edits for real; don't simulate. A hands-on comparison is worth far more than a thought experiment.
- After each experiment, verify the working tree is clean with `git status --short` before moving on.
- Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes
about the repo. Do not read documentation files either. Explore as if you've
never seen it, focusing on code.
- Use git as your safety net — experiment freely. Any edit can be reverted with
`git checkout -- <file>` or `git stash`. Run edits for real; don't simulate. A
hands-on comparison is worth far more than a thought experiment.
- After each experiment, verify the working tree is clean with
`git status --short` before moving on.
### How to compare — correct use only
- **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed for, called the way a competent user
would call it. A tool doing exactly what its contract says is not a finding, even if a careless caller could misuse
it.
- **Know the contract before you call.** Before invoking any tool, have a one-sentence understanding of what it does. If
you expect an error or "not applicable," don't make the call.
- **Refactoring semantics are real.** Inlining requires a substitutable function (typically single-expression, no side
effects); moving requires a legal target; safe-delete requires no surviving usages. If the repo has no suitable
candidate for a given refactoring, report "no suitable candidate in this codebase" and skip it — don't contrive a
broken input.
- **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed
for, called the way a competent user would call it. A tool doing exactly what
its contract says is not a finding, even if a careless caller could misuse it.
- **Know the contract before you call.** Before invoking any tool, have a
one-sentence understanding of what it does. If you expect an error or "not
applicable," don't make the call.
- **Refactoring semantics are real.** Inlining requires a substitutable function
(typically single-expression, no side effects); moving requires a legal
target; safe-delete requires no surviving usages. If the repo has no suitable
candidate for a given refactoring, report "no suitable candidate in this
codebase" and skip it — don't contrive a broken input.
### How to compare — workflow level, not single-call level
- For every task, write out the full end-to-end call chain on each side before drawing conclusions. Include prerequisite
reads and follow-up steps.
- Do not evaluate a tool based on criteria that only arise from mixing workflows incorrectly.
- Ephemeral addressing is a liability. Line numbers and byte offsets go stale after edits; stable addressing (name
paths) may reduce rework.
- For every task, write out the full end-to-end call chain on each side before
drawing conclusions. Include prerequisite reads and follow-up steps.
- Do not evaluate a tool based on criteria that only arise from mixing workflows
incorrectly.
- Ephemeral addressing is a liability. Line numbers and byte offsets go stale
after edits; stable addressing (name paths) may reduce rework.
### How to measure
- Track observations during execution. For every tool call, note: number of calls, approximate input size, output size,
and any prerequisite or verification steps.
- Separate call count, input payload, output payload, and verification cost as distinct axes.
- Track observations during execution. For every tool call, note: number of
calls, approximate input size, output size, and any prerequisite or
verification steps.
- Separate call count, input payload, output payload, and verification cost as
distinct axes.
- Include prerequisite Reads and post-hoc verification steps in comparisons.
When a task falls entirely outside Serena's design scope (e.g., reading config files, small text edits where Edit
already sends minimal payload), classify it as "not applicable" rather than as a negative delta. A negative delta
requires that Serena targets the task and performs worse, not that a tool designed for something else is suboptimal when
misapplied to it.
When a task falls entirely outside Serena's design scope (e.g., reading config
files, small text edits where Edit already sends minimal payload), classify it
as "not applicable" rather than as a negative delta. A negative delta requires
that Serena targets the task and performs worse, not that a tool designed for
something else is suboptimal when misapplied to it.
---
## Exploration phase — tasks to actually perform
Work through the following. Each item exercises a specific capability under correct use; substitute an equivalent if an
item isn't applicable.
Work through the following. Each item exercises a specific capability under
correct use; substitute an equivalent if an item isn't applicable.
### Codebase understanding
1. Get a high-level overview of the repository structure — top-level layout, main packages, entry points.
2. Pick one large source file (300+ lines). Get a structural overview of it. Do it with semantic overview tools and with
Glob/Grep/Read. Then write out the concrete next step on each side and compare the pair of calls, not just the
1. Get a high-level overview of the repository structure — top-level layout,
main packages, entry points.
2. Pick one large source file (300+ lines). Get a structural overview of it. Do
it with semantic overview tools and with Glob/Grep/Read. Then write out the
concrete next step on each side and compare the pair of calls, not just the
overview call.
3. Pick a specific method inside a class and retrieve its body without reading the surrounding file.
4. For one non-trivial symbol, find all references across the codebase. Compare recall and precision under the
question "who uses this in code?" vs "where is this mentioned anywhere, including docs?"
5. For a class, list its subclasses / implementations and its supertypes, including transitively. Compare against what
text search would need to do.
6. For at least one symbol from an external dependency (a third-party library), try to retrieve its definition or
signature. Note whether each toolset can do this at all and what infrastructure it requires (environment activation,
3. Pick a specific method inside a class and retrieve its body without reading
the surrounding file.
4. For one non-trivial symbol, find all references across the codebase. Compare
recall and precision under the question "who uses this in code?" vs "where is
this mentioned anywhere, including docs?"
5. For a class, list its subclasses / implementations and its supertypes,
including transitively. Compare against what text search would need to do.
6. For at least one symbol from an external dependency (a third-party library),
try to retrieve its definition or signature. Note whether each toolset can do
this at all and what infrastructure it requires (environment activation,
site-packages discovery, language-server indexing, etc.).
### Single-file edits — span the full range of edit sizes
7a. Small tweak (1–3 lines inside a method). Change an error message or rename a local variable inside a larger method.
Do it with `Edit` and with symbolic body replacement. Compare payload sent, payload received, and prerequisite reads.
7a. Small tweak (1–3 lines inside a method). Change an error message or rename a
local variable inside a larger method. Do it with `Edit` and with symbolic body
replacement. Compare payload sent, payload received, and prerequisite reads.
7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the main logic of a method while keeping its
signature. Do it both ways.
7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the
main logic of a method while keeping its signature. Do it both ways.
7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire body. Do it both ways.
7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire
body. Do it both ways.
8. Insert a new function/method at a specific structural location (for example, right after an existing method). Try
both the symbolic-insert path and the manual Edit path.
9. Rename a private helper used only within one file. Compare doing it by hand vs. using a semantic rename.
8. Insert a new function/method at a specific structural location (for example,
right after an existing method). Try both the symbolic-insert path and the
manual Edit path.
9. Rename a private helper used only within one file. Compare doing it by hand
vs. using a semantic rename.
### Multi-file changes
10. Rename a symbol (function, class, or method) used across several files including imports. Compare the semantic path
against the built-in equivalent chain.
11. Move a symbol from one module to another, updating imports at all call sites. Use the semantic move tool if
available; plan the built-in equivalent honestly.
12. Move a file or package to a different location, updating imports at all call sites. Use the semantic move tool if
available; plan the built-in equivalent honestly.
12. Delete a symbol safely, checking it has no remaining usages. Compare search-then-delete with a safe-delete tool.
13. Delete a symbol and propagate the deletion to all call sites. Compare to how the built-in equivalent would work.
13. Inline a small helper into its call sites — only if the codebase contains a function that is legally inlinable. If
no such candidate exists, report "no suitable candidate" and skip it.
10. Rename a symbol (function, class, or method) used across several files
including imports. Compare the semantic path against the built-in equivalent
chain.
11. Move a symbol from one module to another, updating imports at all call
sites. Use the semantic move tool if available; plan the built-in equivalent
honestly.
12. Move a file or package to a different location, updating imports at all call
sites. Use the semantic move tool if available; plan the built-in equivalent
honestly.
12. Delete a symbol safely, checking it has no remaining usages. Compare
search-then-delete with a safe-delete tool.
13. Delete a symbol and propagate the deletion to all call sites. Compare to how
the built-in equivalent would work.
13. Inline a small helper into its call sites — only if the codebase contains a
function that is legally inlinable. If no such candidate exists, report "no
suitable candidate" and skip it.
### Reliability & correctness under correct use
14. Scope precision. Demonstrate that semantic tools address symbols by name path and can target a specific class
method, override, or overload that text search would over-match.
15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are updated or none. A chain of `Edit`
calls is not.
16. Success signals. For each completed refactor, note what each tool returns on success.
14. Scope precision. Demonstrate that semantic tools address symbols by name
path and can target a specific class method, override, or overload that text
search would over-match.
15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are
updated or none. A chain of `Edit` calls is not.
16. Success signals. For each completed refactor, note what each tool returns on
success.
### Workflow effects across multiple edits
17. Chain at least three edits in one file. Report what each toolset requires between edits.
18. Multi-step exploration across the repo. Note whether intermediate results remain useful across later edits or have
to be refreshed.
17. Chain at least three edits in one file. Report what each toolset requires
between edits.
18. Multi-step exploration across the repo. Note whether intermediate results
remain useful across later edits or have to be refreshed.
### Things where the comparison shouldn't be interesting
19. Read and understand a non-code file (config, changelog, docs, notebook). Semantic-code tools don't apply — use
`Read`.
20. Search for a free-text pattern across the repo (log string, magic constant, URL). Use `Grep`.
19. Read and understand a non-code file (config, changelog, docs, notebook).
Semantic-code tools don't apply — use `Read`.
20. Search for a free-text pattern across the repo (log string, magic constant,
URL). Use `Grep`.
---
## Evaluation phase
Write a report structured for progressive disclosure.
**Value-weighting is required.** For every contribution or difference you identify — positive, neutral, or negative —
estimate:
**Value-weighting is required.** For every contribution or difference you
identify — positive, neutral, or negative — estimate:
- **Frequency:** how often this arises in typical coding work
- **Value per hit:** calls saved, tokens saved, or correctness impact
Order findings by **frequency × value-per-hit**, not novelty.
**Every section must end with a one-sentence verdict** summarizing the practical takeaway.
**Every section must end with a one-sentence verdict** summarizing the practical
takeaway.
---
### 1. Headline: what Serena changes
Open with a precise description of the delta Serena provides.
Distinguish between three categories:
(a) tasks where Serena adds capability,
(b) tasks where Serena applies but offers
no improvement, and (c) tasks outside Serena's scope.
Only category (b) constitutes a neutral or negative finding.
Category (c) is context, not a finding.
Open with a precise description of the delta Serena provides. Distinguish
between three categories: (a) tasks where Serena adds capability, (b) tasks
where Serena applies but offers no improvement, and (c) tasks outside Serena's
scope. Only category (b) constitutes a neutral or negative finding. Category (c)
is context, not a finding.
A reader stopping here should understand both what is gained and what is not.
**Verdict:** (one sentence)
---
### 2. Added value and differences by area (3–6 bullets)
@@ -190,7 +226,7 @@ Avoid framing in terms of “wins”; describe concrete differences.
**Verdict:** (one sentence)
---
### 3. Detailed evidence, grouped by capability
@@ -208,7 +244,7 @@ Include cases where:
End each subsection with a verdict.
---
### 4. Token-efficiency analysis
@@ -222,7 +258,7 @@ Include cases where each toolset is more efficient.
**Verdict:** (one sentence)
---
### 5. Reliability & correctness (under correct use)
@@ -238,35 +274,36 @@ Include both strengths and limitations of each toolset.
**Verdict:** (one sentence)
---
### 6. Workflow effects across a session
Evaluate multi-step workflows and whether advantages compound or diminish. Include neutral or negative findings where
applicable.
Evaluate multi-step workflows and whether advantages compound or diminish.
Include neutral or negative findings where applicable.
**Verdict:** (one sentence)
---
### 7. Unique capabilities (if any)
List capabilities that have no practical built-in equivalent. If none exist, explicitly state that. Annotate each with
frequency and impact.
List capabilities that have no practical built-in equivalent. If none exist,
explicitly state that. Annotate each with frequency and impact.
**Verdict:** (one sentence)
---
### 8. Tasks outside Serena's scope (built-in only)
Identify tasks where built-ins are the natural choice because Serena's tools don't target them. List these briefly for
completeness but do not frame them as Serena shortcomings — they are outside its scope. Estimate their share of daily
Identify tasks where built-ins are the natural choice because Serena's tools
don't target them. List these briefly for completeness but do not frame them as
Serena shortcomings — they are outside its scope. Estimate their share of daily
work to contextualize how much of a session Serena's augmentation covers.
**Verdict:** (one sentence)
---
### 9. Practical usage rule
@@ -274,7 +311,7 @@ Provide a decision rule for choosing between toolsets based on task type.
**Verdict:** (one sentence)
---
## What I'm looking for
@@ -283,7 +320,7 @@ Provide a decision rule for choosing between toolsets based on task type.
- Clear quantification of impact
- Honest workflow-level comparisons
---
## What I am not looking for
@@ -292,4 +329,5 @@ Provide a decision rule for choosing between toolsets based on task type.
- Neutral statements without evidence
- Binary recommendations
- Novelty-driven ordering
- Unquantified claims
- Unquantified claims
```
@@ -0,0 +1,12 @@
# Summary Prompt
We used the prompt below to summarize the evaluation of Serena.
```
Write a one-sentence user-facing summary about the value that Serena's tools
provide for coding. With a bit of emotion but grounded in the evaluation that
you have done above. It should be written from the perspective of a coding AI
agent and make it clear whether an AI agent would ask its owner to add Serena's
tools or not. It will be the first sentence that a potential new user will see
about how useful Serena is for an AI agent.
```
@@ -0,0 +1,17 @@
# Results
This section presents the results of the evaluation.
We performed evaluations using popular AI coding agents in representative scenarios — different
agents, different programming languages, and different codebases — to show that the results are
not specific to a single setup.
All evaluations were conducted using the **JetBrains-powered version** of Serena, as it is the
more powerful backend with a broader set of refactoring and navigation capabilities. The
evaluation can easily be repeated with the LSP-based backend to assess its subset of capabilities.
- [Claude Code (Opus 4.6) on a large Python codebase](010_cc_on_tianshou)
- [Codex (GPT 5.4) on a Java codebase](020_codex_on_jbplugin)
You can run your own evaluation on a project of your choice by reusing our
[evaluation prompt](../020_prompts/010_evaluation-prompt.md).
@@ -1,9 +1,13 @@
# Evaluation Report: Serena's Tools vs Built-In Tools
:::{admonition} Evaluation Result
:class: note
**Generated by**: Claude Opus 4.6 (coding AI agent in Claude Code CLI)
**Codebase:** [Tianshou](https://github.com/thu-ml/tianshou) — a Python reinforcement learning library (~26K lines, 43 source files)
:::
# Claude Code (Opus)
> **One-line summary:** Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up.
**Evaluated by:** Claude Opus 4.6 (coding AI agent in Claude Code CLI)
**Codebase:** [tianshou](https://github.com/thu-ml/tianshou) — a Python reinforcement learning library (~26K lines, 43 source files)
**Method:** Hands-on, side-by-side execution of 20 tasks using both toolsets. All edits were applied to real files and verified via `git diff`, then reverted.
---
@@ -1,14 +1,19 @@
# Evaluation: What Serena Adds Over Built-Ins
:::{admonition} Evaluation Result
:class: note
**Generated by:** GPT-5.4 (high) in Codex
**Codebase:** Serena JetBrains Plugin (Java)
:::
**Evaluated by:** Gpt 5.4 (high) in Codex
# Codex (GPT-5.4)
> **One-line summary:** As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and
refactoring, which makes real code changes feel faster, safer, and far less blind.
This report compares Serena's JetBrains-backed semantic coding tools with built-in file, shell, search, and patch tools in this repository. The comparison assumes competent use of both toolsets: built-ins are used for text, file, shell, config, and small patch work; Serena is used where symbol identity, language semantics, or IDE refactoring semantics apply.
Method: I explored code first, avoided repo documentation and prior notes, ran real edits/refactors, and reverted after each experiment. After each edit/refactor experiment I checked `git status --short` and returned the tree to clean before moving on. Measurements are approximate, but call counts, diff sizes, and result shapes are from observed runs.
> **One-line summary:** As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and
refactoring, which makes real code changes feel faster, safer, and far less blind.
## 1. Headline: what Serena changes
Serena adds a semantic layer over the codebase. Its concrete delta is the ability to address and transform code by symbols: name paths, overload indexes, reference graphs, type hierarchies, external declarations, and JetBrains refactoring operations.
-3
View File
@@ -1,3 +0,0 @@
# Results
Here are the results of the evaluation for various coding agents running on various projects.
+1 -1
View File
@@ -102,7 +102,7 @@ sphinx:
local_extensions : # A list of local extensions to load by sphinx specified by "name: path" items
recursive_update : false # A boolean indicating whether to overwrite the Sphinx config (true) or recursively update (false)
config : # key-value pairs to directly over-ride the Sphinx configuration
master_doc: "01-about/000_intro.md"
master_doc: "01-about/000_evaluation-results.md"
html_theme_options:
logo:
image_light: ../resources/serena-logo.svg
+1 -1
View File
@@ -190,7 +190,7 @@ def autogen_about_intro_features():
autogen_info = f"<!-- This section is auto-generated by {__file__} from the root README.md; do not edit. -->\n\n"
with open(Path(__file__).parent / "01-about" / "000_intro.md", "w", encoding="utf-8") as f:
with open(Path(__file__).parent / "01-about" / "000_evaluation-results.md", "w", encoding="utf-8") as f:
f.write(autogen_info)
f.write("# About Serena\n\n")
f.write(f"**{tagline}**\n\n")