mirror of
https://github.com/tiennm99/serena.git
synced 2026-10-11 03:13:51 +00:00
Restructure evaluation docs
This commit is contained in:
1 parent
c3a5652fc7
commit
46be9c4bd7
12 files changed
+258
-197
No files matched your search
@@ -0,0 +1,2 @@
|
||||
# Evaluation
|
||||
|
||||
@@ -1,46 +1,26 @@
|
||||
# Evaluation
|
||||
# Methodology
|
||||
|
||||
In this section we describe how we evaluate the performance of Serena's tools.
|
||||
In this section we describe the methodology we applied in evaluating the performance of Serena's tools.
|
||||
|
||||
The evaluation measures the **concrete delta** that Serena's tools provide on top of an agent's
|
||||
built-in capabilities (file reads, text edits, grep, shell, etc.).
|
||||
Rather than a simple thumbs-up/thumbs-down, it produces a detailed, evidence-based report
|
||||
covering capabilities, efficiency and reliability.
|
||||
|
||||
## Why Not Benchmarks?
|
||||
|
||||
Standard coding benchmarks (SWE-bench, HumanEval, etc.) measure an agent's ability to solve
|
||||
predefined tasks with a known correct answer. They are valuable for comparing models and agents,
|
||||
but they are a poor fit for evaluating a **tool augmentation layer** like Serena for several reasons:
|
||||
|
||||
- **Benchmarks don't reflect real usage patterns.** Benchmark tasks are typically small, self-contained
|
||||
problems that can be solved by reading and editing a handful of files. They rarely exercise the
|
||||
workflows where Serena's tools shine — cross-file refactoring, navigating large codebases by symbol
|
||||
structure, chaining multiple edits with stable addressing, or querying type hierarchies and
|
||||
external dependencies. A benchmark score would mostly measure performance on tasks where Serena
|
||||
is not expected to help.
|
||||
- **Results would not generalise to the user's project.** Serena's value depends on the codebase
|
||||
(size, language, complexity), the agent (model, built-in tools), and the client harness
|
||||
(Claude Code, Codex, IDE plugins, etc.). A fixed benchmark on a fixed codebase with a fixed
|
||||
agent tells you little about what Serena would add to *your* setup.
|
||||
- **Predefined tasks bias the measurement.** Choosing specific tasks to evaluate inevitably
|
||||
introduces selection bias — we would end up picking tasks that either favour or disfavour Serena.
|
||||
We wanted an evaluation that systematically covers the full surface area of Serena's capabilities
|
||||
without cherry-picking.
|
||||
|
||||
## Design Goals
|
||||
|
||||
Instead of benchmarks, we designed an evaluation methodology with three goals:
|
||||
We designed an evaluation methodology with three goals:
|
||||
|
||||
1. **Reproducible by any user, on any project, with any agent.**
|
||||
1. **Generality.** The evaluation mechanism should be broadly applicable and repeatable by any user, on any project, with any agent.
|
||||
The evaluation is a single prompt that you give to your agent of choice, pointed at your codebase
|
||||
of choice, in your client of choice. This means the results directly reflect the value Serena
|
||||
would add to your actual workflow — not to an artificial benchmark setup. There is nothing to
|
||||
install, configure, or script beyond what you already have.
|
||||
|
||||
2. **The agent evaluates itself.**
|
||||
We deliberately let the AI agent be both the executor and the evaluator. This may seem
|
||||
counterintuitive, but it is the right design choice: the agent is the actual end user of the
|
||||
We deliberately let the AI agent be both the executor and the evaluator.
|
||||
Since most evaluations are strictly quantitative, this may seem unusual, but it is a reasonable choice:
|
||||
The agent is the actual end user of the
|
||||
tools, so it is in the best position to judge whether a semantic tool improves its workflow
|
||||
compared to its built-in alternatives. It can measure call counts, payload sizes, and
|
||||
prerequisite steps from direct experience rather than from proxy metrics. It also avoids the
|
||||
@@ -48,7 +28,7 @@ Instead of benchmarks, we designed an evaluation methodology with three goals:
|
||||
simply uses them and reports what it observes.
|
||||
|
||||
3. **Comprehensive and unbiased by design.**
|
||||
Rather than selecting specific tasks, the prompt defines **task categories** that systematically
|
||||
Rather than selecting specific tasks, the prompt defines *task categories* that systematically
|
||||
span Serena's capabilities: codebase understanding, single-file edits of varying sizes, multi-file
|
||||
refactoring, reliability properties, and workflow effects. The agent picks concrete instances
|
||||
from the codebase at hand, performs each task using both toolsets side by side, and classifies
|
||||
@@ -58,7 +38,7 @@ Instead of benchmarks, we designed an evaluation methodology with three goals:
|
||||
|
||||
## Method
|
||||
|
||||
We give an AI coding agent a single, detailed [evaluation prompt](010_evaluation-prompt) in a one-shot session.
|
||||
We give an AI coding agent a single, detailed [evaluation prompt](020_prompts/010_evaluation-prompt.md) in a one-shot session.
|
||||
The prompt instructs the agent to perform approximately 20 hands-on tasks across five areas:
|
||||
|
||||
1. **Codebase understanding** — structural overviews, targeted symbol retrieval, reference finding, type hierarchies, and external dependency lookup.
|
||||
@@ -72,30 +52,16 @@ tools and its own built-in tools), applies real edits verified via `git diff`, a
|
||||
payload sizes and prerequisite steps. Edits are reverted after each experiment to keep the working tree clean.
|
||||
|
||||
The resulting report classifies each finding into one of three categories:
|
||||
**(a)** tasks where Serena adds capability,
|
||||
**(b)** tasks where Serena applies but offers no improvement, and
|
||||
**(c)** tasks outside Serena's scope.
|
||||
|
||||
* **(a)** tasks where Serena adds capability,
|
||||
* **(b)** tasks where Serena applies but offers no improvement, and
|
||||
* **(c)** tasks outside Serena's scope.
|
||||
|
||||
Only category (b) constitutes a neutral or negative finding; category (c) is context, not a finding.
|
||||
|
||||
After the evaluation, a separate [follow-up prompt](011_followup-summary-prompt) asks the agent for a
|
||||
After the evaluation, a separate [follow-up prompt](020_prompts/020_summary-prompt.md) asks the agent for a
|
||||
one-sentence, user-facing recommendation — the quotes shown on the [main page](https://github.com/oraios/serena).
|
||||
|
||||
## Results
|
||||
|
||||
We performed evaluations using popular AI coding agents in representative scenarios — different
|
||||
agents, different programming languages, and different codebases — to show that the results are
|
||||
not specific to a single setup.
|
||||
|
||||
All evaluations were conducted using the **JetBrains-powered version** of Serena, as it is the
|
||||
more powerful backend with a broader set of refactoring and navigation capabilities. The
|
||||
evaluation can easily be repeated with the LSP-based backend to assess its subset of capabilities.
|
||||
|
||||
- [Claude Code (Opus 4.6) on a large Python codebase](results/010_cc_on_tianshou-serena-evaluation) — tianshou, a reinforcement learning library (~26K lines).
|
||||
- [Codex (GPT 5.4) on a Java codebase](results/020_codex_on_jbplugin-serena-evaluation) — the Serena JetBrains plugin itself.
|
||||
|
||||
You can run your own evaluation on a project of your choice by reusing our
|
||||
[evaluation prompt](010_evaluation-prompt) or adapting it to your needs.
|
||||
|
||||
## Assessment of the Methodology
|
||||
|
||||
_The following assessment was written by Claude Opus 4.6 (high effort) after reading the full
|
||||
@@ -144,14 +110,12 @@ One thing to watch for is whether agents that are less capable than Opus 4.6 or
|
||||
meaningful evaluations — the prompt is quite demanding. But that is a question about the agent, not
|
||||
about the method.
|
||||
|
||||
## Prompt Fairness
|
||||
### Prompt Fairness
|
||||
|
||||
We put significant effort in crafting an unbiased, fair and objective evaluation prompt.
|
||||
The only "biased" aspects are some sentences about misuse of tools, which we consider irrelevant for the purpose of evaluation
|
||||
(sufficiently intelligent models use the tools correctly).
|
||||
Here is what Opus 4.6 thinks about our prompt:
|
||||
We put significant effort into crafting an unbiased, fair evaluation prompt.
|
||||
To assess the prompt's fairness, we asked Claude Opus 4.6 to evaluate the prompt itself.
|
||||
|
||||
**Q**:
|
||||
**Prompt**:
|
||||
> The attached prompt is designed to evaluate the additive value of Serena as an augmentation layer on top of built-in tools
|
||||
> (not to compare two competing alternatives).
|
||||
> Is the methodology suitable for that goal, or does it introduce bias that would distort the measurement of what Serena adds?
|
||||
@@ -160,12 +124,37 @@ Here is what Opus 4.6 thinks about our prompt:
|
||||
> Do not question the validity of this assumption in your evaluation of the prompt.
|
||||
> Give a brief answer, maximally two sentences.
|
||||
>
|
||||
> \<Evalution prompt pasted here\>
|
||||
> \<evaluation prompt pasted here\>
|
||||
|
||||
**Opus' Answer**:
|
||||
**Claude Opus' Answer**:
|
||||
> "The methodology is well-suited to its stated goal: by explicitly scoping to correct use, treating Serena as an
|
||||
> augmentation layer rather than a competitor, and classifying out-of-scope tasks as context rather than negative
|
||||
> findings, it avoids the most common biases that would distort a delta measurement. The one risk is that the "correct-use
|
||||
> rule" could subtly favor Serena by preemptively filtering out any scenario where its tools underperform — but the prompt
|
||||
> mitigates this by requiring category (b) findings (applies but no improvement) and explicitly demanding negative deltas
|
||||
> be reported, so the bias is acknowledged and structurally counterbalanced."
|
||||
|
||||
According to Claude, the only "biased" aspects are some sentences about misuse of tools, which we consider irrelevant for the purpose of evaluation
|
||||
(sufficiently intelligent models use the tools correctly).
|
||||
|
||||
|
||||
## Why Not Benchmarks?
|
||||
|
||||
Standard coding benchmarks (SWE-bench, HumanEval, etc.) measure an agent's ability to solve
|
||||
predefined tasks with a known correct answer. They are valuable for comparing models and agents,
|
||||
but they are a poor fit for evaluating a **tool augmentation layer** like Serena for several reasons:
|
||||
|
||||
- **Benchmarks don't reflect real usage patterns.** Benchmark tasks are typically small, self-contained
|
||||
problems that can be solved by reading and editing a handful of files. They rarely exercise the
|
||||
workflows where Serena's tools shine — cross-file refactoring, navigating large codebases by symbol
|
||||
structure, chaining multiple edits with stable addressing, or querying type hierarchies and
|
||||
external dependencies. A benchmark score would mostly measure performance on tasks where Serena
|
||||
is not expected to help.
|
||||
- **Results would not generalise to the user's project.** Serena's value depends on the codebase
|
||||
(size, language, complexity), the agent (model, built-in tools), and the client harness
|
||||
(Claude Code, Codex, IDE plugins, etc.). A fixed benchmark on a fixed codebase with a fixed
|
||||
agent tells you little about what Serena would add to *your* setup.
|
||||
- **Predefined tasks bias the measurement.** Choosing specific tasks to evaluate inevitably
|
||||
introduces selection bias — we would end up picking tasks that either favour or disfavour Serena.
|
||||
We wanted an evaluation that systematically covers the full surface area of Serena's capabilities
|
||||
without cherry-picking.
|
||||
@@ -1,9 +0,0 @@
|
||||
# Post-Evaluation Summary Prompt
|
||||
|
||||
_Use the prompt below to summarize the evaluation of Serena by your agent.
|
||||
We used it to create the recommendations we display in the readme._
|
||||
|
||||
Write a one-sentence user-facing summary about the value that Serena's tools provide for coding. With a bit
|
||||
of emotion but grounded in the evaluation that you have done above. It should be written from the perspective of a coding AI agent and make it
|
||||
clear whether an AI agent would ask its owner to add Serena's tools or not. It will be the first sentence that a potential new user will see
|
||||
about how useful Serena is for an AI agent.
|
||||
@@ -0,0 +1,6 @@
|
||||
# Prompts
|
||||
|
||||
Our evaluation uses two prompts, which are passed to the LLM in order to generate the evaluation results:
|
||||
|
||||
1. the [main evaluation prompt](010_evaluation-prompt), which performs the analysis
|
||||
2. the [summary prompt](020_summary-prompt), which condenses the results
|
||||
+156
-118
@@ -1,182 +1,218 @@
|
||||
# Evaluation Prompt
|
||||
|
||||
Use the prompt below to evaluate the added value of Serena's tools against your agent's built-in tools on a project of
|
||||
your choice.
|
||||
All evaluations that you find in our documentation were created in one-shot sessions, only using this prompt and
|
||||
then following up with a separate [summary prompt](011_followup-summary-prompt)
|
||||
We use the prompt below to evaluate the added value of Serena's tools against the
|
||||
agent's built-in tools on a given project.
|
||||
The evaluations were created in one-shot sessions, only using this
|
||||
prompt and the follow-up [prompt for summarization](020_summary-prompt)
|
||||
|
||||
```
|
||||
# Evaluate Serena's Tools Against Built-Ins
|
||||
|
||||
You have access to Serena's coding tools alongside your built-in tools (Read, Edit, Write, Glob, Grep, Bash, etc.). I
|
||||
want a thorough, evidence-based evaluation of **what Serena's tools add on top of the built-ins**, assuming both
|
||||
You have access to Serena's coding tools alongside your built-in tools (Read,
|
||||
Edit, Write, Glob, Grep, Bash, etc.). I want a thorough, evidence-based
|
||||
evaluation of **what Serena's tools add on top of the built-ins**, assuming both
|
||||
toolsets are used correctly.
|
||||
|
||||
This is an evaluation, not a user guide, and it is not a binary adoption pitch. Your job is to answer: *if a competent
|
||||
user of both toolsets had only the built-ins, what concrete capabilities and efficiency differences would they
|
||||
experience, and by how much?* A reader who finishes your report should have a clear, specific picture of what Serena
|
||||
adds — as well as where it provides no meaningful improvement or introduces tradeoffs — in terms of capabilities,
|
||||
workflows, and efficiency. Not a thumbs-up/thumbs-down, but a sharp description of the delta.
|
||||
This is an evaluation, not a user guide, and it is not a binary adoption pitch.
|
||||
Your job is to answer: *if a competent user of both toolsets had only the
|
||||
built-ins, what concrete capabilities and efficiency differences would they
|
||||
experience, and by how much?* A reader who finishes your report should have a
|
||||
clear, specific picture of what Serena adds — as well as where it provides no
|
||||
meaningful improvement or introduces tradeoffs — in terms of capabilities,
|
||||
workflows, and efficiency. Not a thumbs-up/thumbs-down, but a sharp description
|
||||
of the delta.
|
||||
|
||||
Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be careful of X" warnings are out of scope.
|
||||
They belong in onboarding material for a developer learning the tools, not in a delta analysis of what the tools add.
|
||||
Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be
|
||||
careful of X" warnings are out of scope. They belong in onboarding material for
|
||||
a developer learning the tools, not in a delta analysis of what the tools add.
|
||||
|
||||
**Describe the measured differences clearly and neutrally.** Avoid generic or non-informative framing such as "both have
|
||||
their place" unless supported by concrete findings. If Serena adds substantial capabilities, name and quantify them. If
|
||||
it adds marginal or no capabilities, say that and show why. If there are regressions or tradeoffs, include them
|
||||
explicitly. The two toolsets are complementary — that's a given, not the answer.
|
||||
Serena is an augmentation layer, not a replacement. Do not penalize it for tasks it was not designed to address —
|
||||
instead, note those tasks as "built-in only" and move on. The evaluation should measure what Serena adds where it
|
||||
applies, not what it fails to add where it doesn't.The answer is a specific list of what
|
||||
Serena contributes (or does not contribute) to a correct-use workflow relative to built-ins.
|
||||
**Describe the measured differences clearly and neutrally.** Avoid generic or
|
||||
non-informative framing such as "both have their place" unless supported by
|
||||
concrete findings. If Serena adds substantial capabilities, name and quantify
|
||||
them. If it adds marginal or no capabilities, say that and show why. If there
|
||||
are regressions or tradeoffs, include them explicitly. The two toolsets are
|
||||
complementary — that's a given, not the answer. Serena is an augmentation layer,
|
||||
not a replacement. Do not penalize it for tasks it was not designed to address —
|
||||
instead, note those tasks as "built-in only" and move on. The evaluation should
|
||||
measure what Serena adds where it applies, not what it fails to add where it
|
||||
doesn't.The answer is a specific list of what Serena contributes (or does not
|
||||
contribute) to a correct-use workflow relative to built-ins.
|
||||
|
||||
Write the report to serena-evaluation.md in the repo root.
|
||||
|
||||
---
|
||||
|
||||
|
||||
## Ground rules
|
||||
|
||||
### Starting conditions
|
||||
|
||||
- Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes about the repo. Do not read
|
||||
documentation files either. Explore as if you've never seen it, focusing on code.
|
||||
- Use git as your safety net — experiment freely. Any edit can be reverted with `git checkout -- <file>` or `git stash`.
|
||||
Run edits for real; don't simulate. A hands-on comparison is worth far more than a thought experiment.
|
||||
- After each experiment, verify the working tree is clean with `git status --short` before moving on.
|
||||
- Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes
|
||||
about the repo. Do not read documentation files either. Explore as if you've
|
||||
never seen it, focusing on code.
|
||||
- Use git as your safety net — experiment freely. Any edit can be reverted with
|
||||
`git checkout -- <file>` or `git stash`. Run edits for real; don't simulate. A
|
||||
hands-on comparison is worth far more than a thought experiment.
|
||||
- After each experiment, verify the working tree is clean with
|
||||
`git status --short` before moving on.
|
||||
|
||||
### How to compare — correct use only
|
||||
|
||||
- **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed for, called the way a competent user
|
||||
would call it. A tool doing exactly what its contract says is not a finding, even if a careless caller could misuse
|
||||
it.
|
||||
- **Know the contract before you call.** Before invoking any tool, have a one-sentence understanding of what it does. If
|
||||
you expect an error or "not applicable," don't make the call.
|
||||
- **Refactoring semantics are real.** Inlining requires a substitutable function (typically single-expression, no side
|
||||
effects); moving requires a legal target; safe-delete requires no surviving usages. If the repo has no suitable
|
||||
candidate for a given refactoring, report "no suitable candidate in this codebase" and skip it — don't contrive a
|
||||
broken input.
|
||||
- **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed
|
||||
for, called the way a competent user would call it. A tool doing exactly what
|
||||
its contract says is not a finding, even if a careless caller could misuse it.
|
||||
- **Know the contract before you call.** Before invoking any tool, have a
|
||||
one-sentence understanding of what it does. If you expect an error or "not
|
||||
applicable," don't make the call.
|
||||
- **Refactoring semantics are real.** Inlining requires a substitutable function
|
||||
(typically single-expression, no side effects); moving requires a legal
|
||||
target; safe-delete requires no surviving usages. If the repo has no suitable
|
||||
candidate for a given refactoring, report "no suitable candidate in this
|
||||
codebase" and skip it — don't contrive a broken input.
|
||||
|
||||
### How to compare — workflow level, not single-call level
|
||||
|
||||
- For every task, write out the full end-to-end call chain on each side before drawing conclusions. Include prerequisite
|
||||
reads and follow-up steps.
|
||||
- Do not evaluate a tool based on criteria that only arise from mixing workflows incorrectly.
|
||||
- Ephemeral addressing is a liability. Line numbers and byte offsets go stale after edits; stable addressing (name
|
||||
paths) may reduce rework.
|
||||
- For every task, write out the full end-to-end call chain on each side before
|
||||
drawing conclusions. Include prerequisite reads and follow-up steps.
|
||||
- Do not evaluate a tool based on criteria that only arise from mixing workflows
|
||||
incorrectly.
|
||||
- Ephemeral addressing is a liability. Line numbers and byte offsets go stale
|
||||
after edits; stable addressing (name paths) may reduce rework.
|
||||
|
||||
### How to measure
|
||||
|
||||
- Track observations during execution. For every tool call, note: number of calls, approximate input size, output size,
|
||||
and any prerequisite or verification steps.
|
||||
- Separate call count, input payload, output payload, and verification cost as distinct axes.
|
||||
- Track observations during execution. For every tool call, note: number of
|
||||
calls, approximate input size, output size, and any prerequisite or
|
||||
verification steps.
|
||||
- Separate call count, input payload, output payload, and verification cost as
|
||||
distinct axes.
|
||||
- Include prerequisite Reads and post-hoc verification steps in comparisons.
|
||||
|
||||
When a task falls entirely outside Serena's design scope (e.g., reading config files, small text edits where Edit
|
||||
already sends minimal payload), classify it as "not applicable" rather than as a negative delta. A negative delta
|
||||
requires that Serena targets the task and performs worse, not that a tool designed for something else is suboptimal when
|
||||
misapplied to it.
|
||||
When a task falls entirely outside Serena's design scope (e.g., reading config
|
||||
files, small text edits where Edit already sends minimal payload), classify it
|
||||
as "not applicable" rather than as a negative delta. A negative delta requires
|
||||
that Serena targets the task and performs worse, not that a tool designed for
|
||||
something else is suboptimal when misapplied to it.
|
||||
|
||||
|
||||
---
|
||||
|
||||
## Exploration phase — tasks to actually perform
|
||||
|
||||
Work through the following. Each item exercises a specific capability under correct use; substitute an equivalent if an
|
||||
item isn't applicable.
|
||||
Work through the following. Each item exercises a specific capability under
|
||||
correct use; substitute an equivalent if an item isn't applicable.
|
||||
|
||||
### Codebase understanding
|
||||
|
||||
1. Get a high-level overview of the repository structure — top-level layout, main packages, entry points.
|
||||
2. Pick one large source file (300+ lines). Get a structural overview of it. Do it with semantic overview tools and with
|
||||
Glob/Grep/Read. Then write out the concrete next step on each side and compare the pair of calls, not just the
|
||||
1. Get a high-level overview of the repository structure — top-level layout,
|
||||
main packages, entry points.
|
||||
2. Pick one large source file (300+ lines). Get a structural overview of it. Do
|
||||
it with semantic overview tools and with Glob/Grep/Read. Then write out the
|
||||
concrete next step on each side and compare the pair of calls, not just the
|
||||
overview call.
|
||||
3. Pick a specific method inside a class and retrieve its body without reading the surrounding file.
|
||||
4. For one non-trivial symbol, find all references across the codebase. Compare recall and precision under the
|
||||
question "who uses this in code?" vs "where is this mentioned anywhere, including docs?"
|
||||
5. For a class, list its subclasses / implementations and its supertypes, including transitively. Compare against what
|
||||
text search would need to do.
|
||||
6. For at least one symbol from an external dependency (a third-party library), try to retrieve its definition or
|
||||
signature. Note whether each toolset can do this at all and what infrastructure it requires (environment activation,
|
||||
3. Pick a specific method inside a class and retrieve its body without reading
|
||||
the surrounding file.
|
||||
4. For one non-trivial symbol, find all references across the codebase. Compare
|
||||
recall and precision under the question "who uses this in code?" vs "where is
|
||||
this mentioned anywhere, including docs?"
|
||||
5. For a class, list its subclasses / implementations and its supertypes,
|
||||
including transitively. Compare against what text search would need to do.
|
||||
6. For at least one symbol from an external dependency (a third-party library),
|
||||
try to retrieve its definition or signature. Note whether each toolset can do
|
||||
this at all and what infrastructure it requires (environment activation,
|
||||
site-packages discovery, language-server indexing, etc.).
|
||||
|
||||
### Single-file edits — span the full range of edit sizes
|
||||
|
||||
7a. Small tweak (1–3 lines inside a method). Change an error message or rename a local variable inside a larger method.
|
||||
Do it with `Edit` and with symbolic body replacement. Compare payload sent, payload received, and prerequisite reads.
|
||||
7a. Small tweak (1–3 lines inside a method). Change an error message or rename a
|
||||
local variable inside a larger method. Do it with `Edit` and with symbolic body
|
||||
replacement. Compare payload sent, payload received, and prerequisite reads.
|
||||
|
||||
7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the main logic of a method while keeping its
|
||||
signature. Do it both ways.
|
||||
7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the
|
||||
main logic of a method while keeping its signature. Do it both ways.
|
||||
|
||||
7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire body. Do it both ways.
|
||||
7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire
|
||||
body. Do it both ways.
|
||||
|
||||
8. Insert a new function/method at a specific structural location (for example, right after an existing method). Try
|
||||
both the symbolic-insert path and the manual Edit path.
|
||||
9. Rename a private helper used only within one file. Compare doing it by hand vs. using a semantic rename.
|
||||
8. Insert a new function/method at a specific structural location (for example,
|
||||
right after an existing method). Try both the symbolic-insert path and the
|
||||
manual Edit path.
|
||||
9. Rename a private helper used only within one file. Compare doing it by hand
|
||||
vs. using a semantic rename.
|
||||
|
||||
### Multi-file changes
|
||||
|
||||
10. Rename a symbol (function, class, or method) used across several files including imports. Compare the semantic path
|
||||
against the built-in equivalent chain.
|
||||
11. Move a symbol from one module to another, updating imports at all call sites. Use the semantic move tool if
|
||||
available; plan the built-in equivalent honestly.
|
||||
12. Move a file or package to a different location, updating imports at all call sites. Use the semantic move tool if
|
||||
available; plan the built-in equivalent honestly.
|
||||
12. Delete a symbol safely, checking it has no remaining usages. Compare search-then-delete with a safe-delete tool.
|
||||
13. Delete a symbol and propagate the deletion to all call sites. Compare to how the built-in equivalent would work.
|
||||
13. Inline a small helper into its call sites — only if the codebase contains a function that is legally inlinable. If
|
||||
no such candidate exists, report "no suitable candidate" and skip it.
|
||||
10. Rename a symbol (function, class, or method) used across several files
|
||||
including imports. Compare the semantic path against the built-in equivalent
|
||||
chain.
|
||||
11. Move a symbol from one module to another, updating imports at all call
|
||||
sites. Use the semantic move tool if available; plan the built-in equivalent
|
||||
honestly.
|
||||
12. Move a file or package to a different location, updating imports at all call
|
||||
sites. Use the semantic move tool if available; plan the built-in equivalent
|
||||
honestly.
|
||||
12. Delete a symbol safely, checking it has no remaining usages. Compare
|
||||
search-then-delete with a safe-delete tool.
|
||||
13. Delete a symbol and propagate the deletion to all call sites. Compare to how
|
||||
the built-in equivalent would work.
|
||||
13. Inline a small helper into its call sites — only if the codebase contains a
|
||||
function that is legally inlinable. If no such candidate exists, report "no
|
||||
suitable candidate" and skip it.
|
||||
|
||||
### Reliability & correctness under correct use
|
||||
|
||||
14. Scope precision. Demonstrate that semantic tools address symbols by name path and can target a specific class
|
||||
method, override, or overload that text search would over-match.
|
||||
15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are updated or none. A chain of `Edit`
|
||||
calls is not.
|
||||
16. Success signals. For each completed refactor, note what each tool returns on success.
|
||||
14. Scope precision. Demonstrate that semantic tools address symbols by name
|
||||
path and can target a specific class method, override, or overload that text
|
||||
search would over-match.
|
||||
15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are
|
||||
updated or none. A chain of `Edit` calls is not.
|
||||
16. Success signals. For each completed refactor, note what each tool returns on
|
||||
success.
|
||||
|
||||
### Workflow effects across multiple edits
|
||||
|
||||
17. Chain at least three edits in one file. Report what each toolset requires between edits.
|
||||
18. Multi-step exploration across the repo. Note whether intermediate results remain useful across later edits or have
|
||||
to be refreshed.
|
||||
17. Chain at least three edits in one file. Report what each toolset requires
|
||||
between edits.
|
||||
18. Multi-step exploration across the repo. Note whether intermediate results
|
||||
remain useful across later edits or have to be refreshed.
|
||||
|
||||
### Things where the comparison shouldn't be interesting
|
||||
|
||||
19. Read and understand a non-code file (config, changelog, docs, notebook). Semantic-code tools don't apply — use
|
||||
`Read`.
|
||||
20. Search for a free-text pattern across the repo (log string, magic constant, URL). Use `Grep`.
|
||||
19. Read and understand a non-code file (config, changelog, docs, notebook).
|
||||
Semantic-code tools don't apply — use `Read`.
|
||||
20. Search for a free-text pattern across the repo (log string, magic constant,
|
||||
URL). Use `Grep`.
|
||||
|
||||
|
||||
---
|
||||
|
||||
## Evaluation phase
|
||||
|
||||
Write a report structured for progressive disclosure.
|
||||
|
||||
**Value-weighting is required.** For every contribution or difference you identify — positive, neutral, or negative —
|
||||
estimate:
|
||||
**Value-weighting is required.** For every contribution or difference you
|
||||
identify — positive, neutral, or negative — estimate:
|
||||
|
||||
- **Frequency:** how often this arises in typical coding work
|
||||
- **Value per hit:** calls saved, tokens saved, or correctness impact
|
||||
|
||||
Order findings by **frequency × value-per-hit**, not novelty.
|
||||
|
||||
**Every section must end with a one-sentence verdict** summarizing the practical takeaway.
|
||||
**Every section must end with a one-sentence verdict** summarizing the practical
|
||||
takeaway.
|
||||
|
||||
|
||||
---
|
||||
|
||||
### 1. Headline: what Serena changes
|
||||
|
||||
Open with a precise description of the delta Serena provides.
|
||||
Distinguish between three categories:
|
||||
(a) tasks where Serena adds capability,
|
||||
(b) tasks where Serena applies but offers
|
||||
no improvement, and (c) tasks outside Serena's scope.
|
||||
Only category (b) constitutes a neutral or negative finding.
|
||||
Category (c) is context, not a finding.
|
||||
Open with a precise description of the delta Serena provides. Distinguish
|
||||
between three categories: (a) tasks where Serena adds capability, (b) tasks
|
||||
where Serena applies but offers no improvement, and (c) tasks outside Serena's
|
||||
scope. Only category (b) constitutes a neutral or negative finding. Category (c)
|
||||
is context, not a finding.
|
||||
|
||||
A reader stopping here should understand both what is gained and what is not.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
|
||||
### 2. Added value and differences by area (3–6 bullets)
|
||||
|
||||
@@ -190,7 +226,7 @@ Avoid framing in terms of “wins”; describe concrete differences.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
|
||||
### 3. Detailed evidence, grouped by capability
|
||||
|
||||
@@ -208,7 +244,7 @@ Include cases where:
|
||||
|
||||
End each subsection with a verdict.
|
||||
|
||||
---
|
||||
|
||||
|
||||
### 4. Token-efficiency analysis
|
||||
|
||||
@@ -222,7 +258,7 @@ Include cases where each toolset is more efficient.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
|
||||
### 5. Reliability & correctness (under correct use)
|
||||
|
||||
@@ -238,35 +274,36 @@ Include both strengths and limitations of each toolset.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
|
||||
### 6. Workflow effects across a session
|
||||
|
||||
Evaluate multi-step workflows and whether advantages compound or diminish. Include neutral or negative findings where
|
||||
applicable.
|
||||
Evaluate multi-step workflows and whether advantages compound or diminish.
|
||||
Include neutral or negative findings where applicable.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
|
||||
### 7. Unique capabilities (if any)
|
||||
|
||||
List capabilities that have no practical built-in equivalent. If none exist, explicitly state that. Annotate each with
|
||||
frequency and impact.
|
||||
List capabilities that have no practical built-in equivalent. If none exist,
|
||||
explicitly state that. Annotate each with frequency and impact.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
|
||||
### 8. Tasks outside Serena's scope (built-in only)
|
||||
|
||||
Identify tasks where built-ins are the natural choice because Serena's tools don't target them. List these briefly for
|
||||
completeness but do not frame them as Serena shortcomings — they are outside its scope. Estimate their share of daily
|
||||
Identify tasks where built-ins are the natural choice because Serena's tools
|
||||
don't target them. List these briefly for completeness but do not frame them as
|
||||
Serena shortcomings — they are outside its scope. Estimate their share of daily
|
||||
work to contextualize how much of a session Serena's augmentation covers.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
|
||||
### 9. Practical usage rule
|
||||
|
||||
@@ -274,7 +311,7 @@ Provide a decision rule for choosing between toolsets based on task type.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
|
||||
## What I'm looking for
|
||||
|
||||
@@ -283,7 +320,7 @@ Provide a decision rule for choosing between toolsets based on task type.
|
||||
- Clear quantification of impact
|
||||
- Honest workflow-level comparisons
|
||||
|
||||
---
|
||||
|
||||
|
||||
## What I am not looking for
|
||||
|
||||
@@ -293,3 +330,4 @@ Provide a decision rule for choosing between toolsets based on task type.
|
||||
- Binary recommendations
|
||||
- Novelty-driven ordering
|
||||
- Unquantified claims
|
||||
```
|
||||
@@ -0,0 +1,12 @@
|
||||
# Summary Prompt
|
||||
|
||||
We used the prompt below to summarize the evaluation of Serena.
|
||||
|
||||
```
|
||||
Write a one-sentence user-facing summary about the value that Serena's tools
|
||||
provide for coding. With a bit of emotion but grounded in the evaluation that
|
||||
you have done above. It should be written from the perspective of a coding AI
|
||||
agent and make it clear whether an AI agent would ask its owner to add Serena's
|
||||
tools or not. It will be the first sentence that a potential new user will see
|
||||
about how useful Serena is for an AI agent.
|
||||
```
|
||||
@@ -0,0 +1,17 @@
|
||||
# Results
|
||||
|
||||
This section presents the results of the evaluation.
|
||||
|
||||
We performed evaluations using popular AI coding agents in representative scenarios — different
|
||||
agents, different programming languages, and different codebases — to show that the results are
|
||||
not specific to a single setup.
|
||||
|
||||
All evaluations were conducted using the **JetBrains-powered version** of Serena, as it is the
|
||||
more powerful backend with a broader set of refactoring and navigation capabilities. The
|
||||
evaluation can easily be repeated with the LSP-based backend to assess its subset of capabilities.
|
||||
|
||||
- [Claude Code (Opus 4.6) on a large Python codebase](010_cc_on_tianshou)
|
||||
- [Codex (GPT 5.4) on a Java codebase](020_codex_on_jbplugin)
|
||||
|
||||
You can run your own evaluation on a project of your choice by reusing our
|
||||
[evaluation prompt](../020_prompts/010_evaluation-prompt.md).
|
||||
+7
-3
@@ -1,9 +1,13 @@
|
||||
# Evaluation Report: Serena's Tools vs Built-In Tools
|
||||
:::{admonition} Evaluation Result
|
||||
:class: note
|
||||
**Generated by**: Claude Opus 4.6 (coding AI agent in Claude Code CLI)
|
||||
**Codebase:** [Tianshou](https://github.com/thu-ml/tianshou) — a Python reinforcement learning library (~26K lines, 43 source files)
|
||||
:::
|
||||
|
||||
# Claude Code (Opus)
|
||||
|
||||
> **One-line summary:** Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up.
|
||||
|
||||
**Evaluated by:** Claude Opus 4.6 (coding AI agent in Claude Code CLI)
|
||||
**Codebase:** [tianshou](https://github.com/thu-ml/tianshou) — a Python reinforcement learning library (~26K lines, 43 source files)
|
||||
**Method:** Hands-on, side-by-side execution of 20 tasks using both toolsets. All edits were applied to real files and verified via `git diff`, then reverted.
|
||||
|
||||
---
|
||||
+10
-5
@@ -1,14 +1,19 @@
|
||||
# Evaluation: What Serena Adds Over Built-Ins
|
||||
:::{admonition} Evaluation Result
|
||||
:class: note
|
||||
**Generated by:** GPT-5.4 (high) in Codex
|
||||
**Codebase:** Serena JetBrains Plugin (Java)
|
||||
:::
|
||||
|
||||
**Evaluated by:** Gpt 5.4 (high) in Codex
|
||||
# Codex (GPT-5.4)
|
||||
|
||||
|
||||
> **One-line summary:** As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and
|
||||
refactoring, which makes real code changes feel faster, safer, and far less blind.
|
||||
|
||||
This report compares Serena's JetBrains-backed semantic coding tools with built-in file, shell, search, and patch tools in this repository. The comparison assumes competent use of both toolsets: built-ins are used for text, file, shell, config, and small patch work; Serena is used where symbol identity, language semantics, or IDE refactoring semantics apply.
|
||||
|
||||
Method: I explored code first, avoided repo documentation and prior notes, ran real edits/refactors, and reverted after each experiment. After each edit/refactor experiment I checked `git status --short` and returned the tree to clean before moving on. Measurements are approximate, but call counts, diff sizes, and result shapes are from observed runs.
|
||||
|
||||
> **One-line summary:** As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and
|
||||
refactoring, which makes real code changes feel faster, safer, and far less blind.
|
||||
|
||||
## 1. Headline: what Serena changes
|
||||
|
||||
Serena adds a semantic layer over the codebase. Its concrete delta is the ability to address and transform code by symbols: name paths, overload indexes, reference graphs, type hierarchies, external declarations, and JetBrains refactoring operations.
|
||||
@@ -1,3 +0,0 @@
|
||||
# Results
|
||||
|
||||
Here are the results of the evaluation for various coding agents running on various projects.
|
||||
+1
-1
@@ -102,7 +102,7 @@ sphinx:
|
||||
local_extensions : # A list of local extensions to load by sphinx specified by "name: path" items
|
||||
recursive_update : false # A boolean indicating whether to overwrite the Sphinx config (true) or recursively update (false)
|
||||
config : # key-value pairs to directly over-ride the Sphinx configuration
|
||||
master_doc: "01-about/000_intro.md"
|
||||
master_doc: "01-about/000_evaluation-results.md"
|
||||
html_theme_options:
|
||||
logo:
|
||||
image_light: ../resources/serena-logo.svg
|
||||
|
||||
@@ -190,7 +190,7 @@ def autogen_about_intro_features():
|
||||
|
||||
autogen_info = f"<!-- This section is auto-generated by {__file__} from the root README.md; do not edit. -->\n\n"
|
||||
|
||||
with open(Path(__file__).parent / "01-about" / "000_intro.md", "w", encoding="utf-8") as f:
|
||||
with open(Path(__file__).parent / "01-about" / "000_evaluation-results.md", "w", encoding="utf-8") as f:
|
||||
f.write(autogen_info)
|
||||
f.write("# About Serena\n\n")
|
||||
f.write(f"**{tagline}**\n\n")
|
||||
|
||||
Reference in new issue
Block a user