mirror of
https://github.com/tiennm99/serena.git
synced 2026-10-11 03:13:51 +00:00
Added documentation on Serena evaluation and first evaluation results
This commit is contained in:
1 parent
e633c94e3b
commit
a3cf1b17c9
6 files changed
+1243
-387
No files matched your search
@@ -30,24 +30,27 @@ more complex codebases.
|
||||
|
||||
## What Our "End Users" Say
|
||||
|
||||
Our end users are essentially AI agents, so they are also in the best position to evaluate Serena.
|
||||
We wrote an unbiased and detailed evaluation prompt, which leads the agent to estimate the value-add of Serena's tools
|
||||
when added on top of their existing capabilities. After about 20 minutes of thorough evaluation and experimentation, here
|
||||
is what the agents had to say:
|
||||
While it is humans who download and set up Serena, our end users are essentially AI agents,
|
||||
so they are also in the best position to evaluate Serena.
|
||||
We crafted an unbiased and detailed prompt which leads the agent to estimate the value of adding Serena's tools
|
||||
to its built-in capabilities. The thorough evaluation usually takes around 25 minutes (per project)
|
||||
and tests every single aspect of Serena. A summary of what the agents had to say:
|
||||
|
||||
**Opus 4.6 in Claude Code, evaluated on a large Python codebase:**
|
||||
> "If I could ask my owner for one upgrade to my coding toolkit, it would be Serena — its symbolic addressing eliminated the constant re-reading that quietly eats my context on multi-edit sessions, its atomic cross-file refactorings collapsed 10-call manual workflows into one, and after a rigorous hands-on evaluation I can say that while built-in tools remain the right choice for small edits and text searches, Serena fills the exact gaps that make me feel clumsy without it."
|
||||
**Opus 4.6 (high effort) in Claude Code on a large Python codebase:**
|
||||
> "Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that
|
||||
would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up."
|
||||
|
||||
**Gpt 5.4 max in Codex CLI, evaluated on a Java codebase:**
|
||||
> "As a coding AI agent, I would ask my owner to add Serena because it lets me work with code as stable symbols instead of fragile text, turning the
|
||||
refactors, lookups, and multi-file edits that usually feel risky and repetitive into precise IDE-backed operations I can trust."
|
||||
**Gpt 5.4 (high) in Codex CLI on a Java codebase:**
|
||||
> "As a coding AI agent, I would ask my owner to add Serena because it gives me the missing IDE-level understanding of symbols, references, and
|
||||
refactorings, turning fragile text surgery into calmer, faster, more confident code changes where semantics matter."
|
||||
|
||||
See our documentation on the [evaluation methodology](https://oraios.github.io/serena/04-evaluation/000_intro.html) and the
|
||||
detailed results beyond the brief recommendations above.
|
||||
You can easily use our methods to run your own evaluation of Serena on a project of your choice.
|
||||
|
||||
Your agent deserves the best coding tools, give them Serena!
|
||||
|
||||
See the documentation on our [evaluation methods](https://oraios.github.io/serena/04-evaluation/000_intro.html) and the
|
||||
detailed results beyond the brief recommendations above. You can easily run your own evaluation of Serena on a project of your choice
|
||||
by reusing our methods or adapting them to your needs.
|
||||
|
||||
|
||||
## How Serena Works
|
||||
|
||||
Serena provides the necessary [tools](https://oraios.github.io/serena/01-about/035_tools.html) for coding workflows,
|
||||
|
||||
@@ -1,3 +1,171 @@
|
||||
# Evaluation
|
||||
|
||||
In this section we describe how we evaluate the performance of Serena's tools.
|
||||
|
||||
The evaluation measures the **concrete delta** that Serena's tools provide on top of an agent's
|
||||
built-in capabilities (file reads, text edits, grep, shell, etc.).
|
||||
Rather than a simple thumbs-up/thumbs-down, it produces a detailed, evidence-based report
|
||||
covering capabilities, efficiency and reliability.
|
||||
|
||||
### Why Not Benchmarks?
|
||||
|
||||
Standard coding benchmarks (SWE-bench, HumanEval, etc.) measure an agent's ability to solve
|
||||
predefined tasks with a known correct answer. They are valuable for comparing models and agents,
|
||||
but they are a poor fit for evaluating a **tool augmentation layer** like Serena for several reasons:
|
||||
|
||||
- **Benchmarks don't reflect real usage patterns.** Benchmark tasks are typically small, self-contained
|
||||
problems that can be solved by reading and editing a handful of files. They rarely exercise the
|
||||
workflows where Serena's tools shine — cross-file refactoring, navigating large codebases by symbol
|
||||
structure, chaining multiple edits with stable addressing, or querying type hierarchies and
|
||||
external dependencies. A benchmark score would mostly measure performance on tasks where Serena
|
||||
is not expected to help.
|
||||
- **Results would not generalise to the user's project.** Serena's value depends on the codebase
|
||||
(size, language, complexity), the agent (model, built-in tools), and the client harness
|
||||
(Claude Code, Codex, IDE plugins, etc.). A fixed benchmark on a fixed codebase with a fixed
|
||||
agent tells you little about what Serena would add to *your* setup.
|
||||
- **Predefined tasks bias the measurement.** Choosing specific tasks to evaluate inevitably
|
||||
introduces selection bias — we would end up picking tasks that either favour or disfavour Serena.
|
||||
We wanted an evaluation that systematically covers the full surface area of Serena's capabilities
|
||||
without cherry-picking.
|
||||
|
||||
### Design Goals
|
||||
|
||||
Instead of benchmarks, we designed an evaluation methodology with three goals:
|
||||
|
||||
1. **Reproducible by any user, on any project, with any agent.**
|
||||
The evaluation is a single prompt that you give to your agent of choice, pointed at your codebase
|
||||
of choice, in your client of choice. This means the results directly reflect the value Serena
|
||||
would add to your actual workflow — not to an artificial benchmark setup. There is nothing to
|
||||
install, configure, or script beyond what you already have.
|
||||
|
||||
2. **The agent evaluates itself.**
|
||||
We deliberately let the AI agent be both the executor and the evaluator. This may seem
|
||||
counterintuitive, but it is the right design choice: the agent is the actual end user of the
|
||||
tools, so it is in the best position to judge whether a semantic tool improves its workflow
|
||||
compared to its built-in alternatives. It can measure call counts, payload sizes, and
|
||||
prerequisite steps from direct experience rather than from proxy metrics. It also avoids the
|
||||
problem of a human evaluator having to simulate how an agent would use the tools — the agent
|
||||
simply uses them and reports what it observes.
|
||||
|
||||
3. **Comprehensive and unbiased by design.**
|
||||
Rather than selecting specific tasks, the prompt defines **task categories** that systematically
|
||||
span Serena's capabilities: codebase understanding, single-file edits of varying sizes, multi-file
|
||||
refactoring, reliability properties, and workflow effects. The agent picks concrete instances
|
||||
from the codebase at hand, performs each task using both toolsets side by side, and classifies
|
||||
every finding as a positive delta, a neutral/negative delta, or out of scope. The prompt
|
||||
explicitly requires reporting negative deltas and cases where Serena offers no improvement,
|
||||
structurally counterbalancing any tendency to favour the tool being evaluated.
|
||||
|
||||
### Method
|
||||
|
||||
We give an AI coding agent a single, detailed [evaluation prompt](010_evaluation-prompt) in a one-shot session.
|
||||
The prompt instructs the agent to perform approximately 20 hands-on tasks across five areas:
|
||||
|
||||
1. **Codebase understanding** — structural overviews, targeted symbol retrieval, reference finding, type hierarchies, and external dependency lookup.
|
||||
2. **Single-file edits** — small tweaks, medium rewrites, full-body replacements, insertions, and local renames.
|
||||
3. **Multi-file changes** — cross-file renames, symbol and file moves, safe deletes, and inlining.
|
||||
4. **Reliability & correctness** — scope precision, atomicity, and success signals.
|
||||
5. **Workflow effects** — chained edits, stable vs. ephemeral addressing, and multi-step exploration.
|
||||
|
||||
For every task, the agent executes the full end-to-end workflow using **both** toolsets (Serena's semantic
|
||||
tools and its own built-in tools), applies real edits verified via `git diff`, and records call counts,
|
||||
payload sizes and prerequisite steps. Edits are reverted after each experiment to keep the working tree clean.
|
||||
|
||||
The resulting report classifies each finding into one of three categories:
|
||||
**(a)** tasks where Serena adds capability,
|
||||
**(b)** tasks where Serena applies but offers no improvement, and
|
||||
**(c)** tasks outside Serena's scope.
|
||||
Only category (b) constitutes a neutral or negative finding; category (c) is context, not a finding.
|
||||
|
||||
After the evaluation, a separate [follow-up prompt](011_followup-summary-prompt) asks the agent for a
|
||||
one-sentence, user-facing recommendation — the quotes shown on the [main page](https://github.com/oraios/serena).
|
||||
|
||||
### Results
|
||||
|
||||
We performed evaluations using popular AI coding agents in representative scenarios — different
|
||||
agents, different programming languages, and different codebases — to show that the results are
|
||||
not specific to a single setup.
|
||||
|
||||
All evaluations were conducted using the **JetBrains-powered version** of Serena, as it is the
|
||||
more powerful backend with a broader set of refactoring and navigation capabilities. The
|
||||
evaluation can easily be repeated with the LSP-based backend to assess its subset of capabilities.
|
||||
|
||||
- [Claude Code (Opus 4.6) on a large Python codebase](results/010_cc_on_tianshou-serena-evaluation) — tianshou, a reinforcement learning library (~26K lines).
|
||||
- [Codex (GPT 5.4) on a Java codebase](results/020_codex_on_jbplugin-serena-evaluation) — the Serena JetBrains plugin itself.
|
||||
|
||||
You can run your own evaluation on a project of your choice by reusing our
|
||||
[evaluation prompt](010_evaluation-prompt) or adapting it to your needs.
|
||||
|
||||
### Assessment of the Methodology
|
||||
|
||||
_The following assessment was written by Claude Opus 4.6 (high effort) after reading the full
|
||||
evaluation methodology, evaluation prompt, and all result documents._
|
||||
|
||||
**The methodology is sound**, and the two published results demonstrate that it produces meaningful,
|
||||
detailed evaluations. Both reports follow the prompt's structure faithfully: they perform all ~20
|
||||
tasks using both toolsets, record concrete measurements (call counts, payload sizes, prerequisite
|
||||
steps), and classify findings into the three required categories. Crucially, both agents report
|
||||
neutral and negative deltas honestly — the Claude Code report explicitly notes that small edits are
|
||||
more efficient with built-ins (~4.5x less payload), and the Codex report flags that tiny intra-method
|
||||
changes and simple one-file renames see no benefit from Serena. Neither report reads as promotional;
|
||||
they read as technical comparisons with quantified evidence on both sides.
|
||||
|
||||
The two reports also validate the methodology's design goal of generalisability: despite being
|
||||
produced by different agents (Opus 4.6 vs GPT 5.4), on different codebases (Python RL library vs
|
||||
Java IDE plugin), they converge on the same core findings — cross-file refactoring is Serena's
|
||||
highest-value contribution, structural navigation provides a moderate advantage, and small local
|
||||
edits are better handled by built-ins. This convergence across independent runs increases confidence
|
||||
that the findings reflect genuine properties of the toolset rather than artefacts of a particular
|
||||
agent or codebase.
|
||||
|
||||
The methodology's core strengths are:
|
||||
|
||||
- **Self-evaluation by the agent is the right design choice.** The agent is the actual consumer of the
|
||||
tools, so it can report on workflow friction, payload overhead, and call counts from first-hand
|
||||
experience. A human evaluator would have to guess at these.
|
||||
- **Task categories instead of fixed tasks** avoid cherry-picking while still ensuring coverage. Letting
|
||||
the agent pick concrete instances from the codebase at hand means the evaluation naturally adapts to
|
||||
what the project actually contains (e.g. skipping inline if no suitable candidate exists, as the
|
||||
Claude Code report did for Python, while the Codex report found a suitable Java candidate).
|
||||
- **The three-category classification** (adds value / applies but no improvement / out of scope) is the
|
||||
right framing for an augmentation layer. It prevents the common trap of penalising a tool for not
|
||||
covering things it was never designed to cover.
|
||||
- **Reproducibility by users** is a strong differentiator. Anyone can validate the claims on their own
|
||||
codebase with their own agent.
|
||||
|
||||
The main limitation is scope: the published results cover two agent/codebase combinations using the
|
||||
JetBrains backend. Single-shot variance means a second run of the same agent could produce different
|
||||
task selections and slightly different verdicts, and the LSP backend — which most users will
|
||||
encounter first — has not yet been evaluated. These are practical constraints of cost, not
|
||||
methodological flaws, and the reproducible design explicitly invites users to fill these gaps
|
||||
themselves.
|
||||
|
||||
One thing to watch for is whether agents that are less capable than Opus 4.6 or GPT 5.4 produce
|
||||
meaningful evaluations — the prompt is quite demanding. But that is a question about the agent, not
|
||||
about the method.
|
||||
|
||||
### Prompt Fairness
|
||||
|
||||
We put significant effort in crafting an unbiased, fair and objective evaluation prompt.
|
||||
The only "biased" aspects are some sentences about misuse of tools, which we consider irrelevant for the purpose of evaluation
|
||||
(sufficiently intelligent models use the tools correctly).
|
||||
Here is what Opus 4.6 thinks about our prompt:
|
||||
|
||||
**Q**:
|
||||
> The attached prompt is designed to evaluate the additive value of Serena as an augmentation layer on top of built-in tools
|
||||
> (not to compare two competing alternatives).
|
||||
> Is the methodology suitable for that goal, or does it introduce bias that would distort the measurement of what Serena adds?
|
||||
> Note that we assume that sufficiently intelligent models use the tools correctly,
|
||||
> which is why misuse and failure modes are out of scope for this evaluation.
|
||||
> Do not question the validity of this assumption in your evaluation of the prompt.
|
||||
> Give a brief answer, maximally two sentences.
|
||||
>
|
||||
> \<Evalution prompt pasted here\>
|
||||
|
||||
**Opus' Answer**:
|
||||
> "The methodology is well-suited to its stated goal: by explicitly scoping to correct use, treating Serena as an
|
||||
> augmentation layer rather than a competitor, and classifying out-of-scope tasks as context rather than negative
|
||||
> findings, it avoids the most common biases that would distort a delta measurement. The one risk is that the "correct-use
|
||||
> rule" could subtly favor Serena by preemptively filtering out any scenario where its tools underperform — but the prompt
|
||||
> mitigates this by requiring category (b) findings (applies but no improvement) and explicitly demanding negative deltas
|
||||
> be reported, so the bias is acknowledged and structurally counterbalanced."
|
||||
@@ -1,149 +1,295 @@
|
||||
# Evaluation Prompt
|
||||
|
||||
Use the prompt below to evaluate the added value of Serena's tools against your agent's built-in tools on a project of your choice.
|
||||
Use the prompt below to evaluate the added value of Serena's tools against your agent's built-in tools on a project of
|
||||
your choice.
|
||||
All evaluations that you find in our documentation were created in one-shot sessions, only using this prompt and
|
||||
then following up with a separate [summary prompt](011_followup-summary-prompt)
|
||||
|
||||
|
||||
# Evaluate Serena's Tools Against Built-Ins
|
||||
|
||||
You have access to Serena's coding tools alongside your built-in tools (Read, Edit, Write, Glob, Grep, Bash, etc.). I want a thorough, evidence-based evaluation of **what Serena's tools add on top of the built-ins**, assuming both toolsets are used correctly.
|
||||
You have access to Serena's coding tools alongside your built-in tools (Read, Edit, Write, Glob, Grep, Bash, etc.). I
|
||||
want a thorough, evidence-based evaluation of **what Serena's tools add on top of the built-ins**, assuming both
|
||||
toolsets are used correctly.
|
||||
|
||||
This is an evaluation, not a user guide, and it is not a binary adoption pitch. Your job is to answer: *if a competent user of both toolsets had only the built-ins, what concrete capabilities and efficiency wins would they be missing, and by how much?* A reader who finishes your report should have a clear, specific picture of what Serena adds — in terms of capabilities that weren't available at all, workflows that collapsed from many calls to one, and efficiency multipliers that show up across a session. Not a thumbs-up/thumbs-down, but a sharp description of the delta.
|
||||
This is an evaluation, not a user guide, and it is not a binary adoption pitch. Your job is to answer: *if a competent
|
||||
user of both toolsets had only the built-ins, what concrete capabilities and efficiency differences would they
|
||||
experience, and by how much?* A reader who finishes your report should have a clear, specific picture of what Serena
|
||||
adds — as well as where it provides no meaningful improvement or introduces tradeoffs — in terms of capabilities,
|
||||
workflows, and efficiency. Not a thumbs-up/thumbs-down, but a sharp description of the delta.
|
||||
|
||||
Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be careful of X" warnings are out of scope. They belong in onboarding material for a developer learning the tools, not in a delta analysis of what the tools add.
|
||||
Failure modes from misuse, silent-failure traps, gotcha comparisons, and "be careful of X" warnings are out of scope.
|
||||
They belong in onboarding material for a developer learning the tools, not in a delta analysis of what the tools add.
|
||||
|
||||
**Describe the added value sharply, don't hedge.** A "both have their place" or "complementary, not rivals" opener is not a description of added value — it conveys zero information about what a reader would actually gain or lose by adding the tool. If Serena adds substantial capabilities, name them and quantify them. If it adds marginal capabilities, say that and show why. The two toolsets are complementary — that's a given, not the answer. The answer is a specific list of what Serena contributes to a correct-use workflow that built-ins alone cannot provide.
|
||||
**Describe the measured differences clearly and neutrally.** Avoid generic or non-informative framing such as "both have
|
||||
their place" unless supported by concrete findings. If Serena adds substantial capabilities, name and quantify them. If
|
||||
it adds marginal or no capabilities, say that and show why. If there are regressions or tradeoffs, include them
|
||||
explicitly. The two toolsets are complementary — that's a given, not the answer.
|
||||
Serena is an augmentation layer, not a replacement. Do not penalize it for tasks it was not designed to address —
|
||||
instead, note those tasks as "built-in only" and move on. The evaluation should measure what Serena adds where it
|
||||
applies, not what it fails to add where it doesn't.The answer is a specific list of what
|
||||
Serena contributes (or does not contribute) to a correct-use workflow relative to built-ins.
|
||||
|
||||
Write the report to serena-evaluation.md in the repo root.
|
||||
|
||||
---
|
||||
|
||||
## Ground rules
|
||||
|
||||
### Starting conditions
|
||||
|
||||
- Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes about the repo. Do not read documentation files either. Explore as if you've never seen it, focusing on code.
|
||||
- Use git as your safety net — experiment freely. Any edit can be reverted with `git checkout -- <file>` or `git stash`. Run edits for real; don't simulate. A hands-on comparison is worth far more than a thought experiment.
|
||||
- Start fresh. Do not read project memories, CLAUDE.md shortcuts, or prior notes about the repo. Do not read
|
||||
documentation files either. Explore as if you've never seen it, focusing on code.
|
||||
- Use git as your safety net — experiment freely. Any edit can be reverted with `git checkout -- <file>` or `git stash`.
|
||||
Run edits for real; don't simulate. A hands-on comparison is worth far more than a thought experiment.
|
||||
- After each experiment, verify the working tree is clean with `git status --short` before moving on.
|
||||
|
||||
### How to compare — correct use only
|
||||
|
||||
- **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed for, called the way a competent user would call it. A tool doing exactly what its contract says is not a finding, even if a careless caller could misuse it. Examples of what *not* to report:
|
||||
- Destructive write tools accepting what you send them — that is the contract of a destructive write, not a silent failure mode.
|
||||
- Addressing schemes that require you to know what you're addressing — that is how addressing works.
|
||||
- Tools refusing out-of-scope input (semantic tools on non-code files, single-expression refactorings on multi-branch functions, safe-delete on symbols with usages) — those are correct refusals.
|
||||
- Mid-session friction that only appears if you mix tool families incorrectly — a caller managing their session properly doesn't hit it.
|
||||
- Transient glitches may appear after `git checkout --` or other file-system mutations. Wait a moment before continuing after doing such a checkout, and in case of failure, retry once. If it then succeeds, the finding is a non-finding.
|
||||
- **Know the contract before you call.** Before invoking any tool, have a one-sentence understanding of what it does. If you expect an error or "not applicable," don't make the call.
|
||||
- **Refactoring semantics are real.** Inlining requires a substitutable function (typically single-expression, no side effects); moving requires a legal target; safe-delete requires no surviving usages. If the repo has no suitable candidate for a given refactoring, report "no suitable candidate in this codebase" and skip it — don't contrive a broken input.
|
||||
- **Correct-use rule.** Evaluate each tool on inputs and tasks it was designed for, called the way a competent user
|
||||
would call it. A tool doing exactly what its contract says is not a finding, even if a careless caller could misuse
|
||||
it.
|
||||
- **Know the contract before you call.** Before invoking any tool, have a one-sentence understanding of what it does. If
|
||||
you expect an error or "not applicable," don't make the call.
|
||||
- **Refactoring semantics are real.** Inlining requires a substitutable function (typically single-expression, no side
|
||||
effects); moving requires a legal target; safe-delete requires no surviving usages. If the repo has no suitable
|
||||
candidate for a given refactoring, report "no suitable candidate in this codebase" and skip it — don't contrive a
|
||||
broken input.
|
||||
|
||||
### How to compare — workflow level, not single-call level
|
||||
|
||||
- For every task, write out the full end-to-end call chain on each side before declaring a winner. Not just "Serena does it in one call" vs "Grep returns line N" — spell out every call you'd make to reach the goal, including the next step after whichever tool you called first. Many apparent wins and losses evaporate once you include the follow-up.
|
||||
- Don't score a tool on criteria that only matter in the other toolset's workflow. If you find yourself penalizing Serena for missing a feature (line numbers, text anchors) or penalizing built-ins for missing a feature (name paths, type hierarchy), check whether that feature is actually needed in the tool's own native follow-up. If it's only needed because you're planning to fall back to the other toolset, you're holding the tool to the wrong workflow.
|
||||
- Ephemeral addressing is a liability. Line numbers and byte offsets go stale the moment a file is edited. Stable addressing (name paths) is an efficiency win, even when the output looks smaller.
|
||||
- For every task, write out the full end-to-end call chain on each side before drawing conclusions. Include prerequisite
|
||||
reads and follow-up steps.
|
||||
- Do not evaluate a tool based on criteria that only arise from mixing workflows incorrectly.
|
||||
- Ephemeral addressing is a liability. Line numbers and byte offsets go stale after edits; stable addressing (name
|
||||
paths) may reduce rework.
|
||||
|
||||
### How to measure
|
||||
|
||||
- Track observations while you work, not at the end. For every tool call, note: number of calls needed, approximate size of what you sent (including re-sent content), size of what you got back, and any prerequisite reads or follow-up verification. These are the raw material for the final report.
|
||||
- Separate call count, input payload, output payload, and verification cost as distinct axes. A tool that halves call count but doubles payload may not be a net win.
|
||||
- When comparing two approaches on the same task, include prerequisite Reads and post-hoc verification Greps in the cost — don't hide them outside the comparison. A cross-file rename via Edit isn't "one Edit call"; it's `grep + read × N + edit × N + verify`.
|
||||
- Track observations during execution. For every tool call, note: number of calls, approximate input size, output size,
|
||||
and any prerequisite or verification steps.
|
||||
- Separate call count, input payload, output payload, and verification cost as distinct axes.
|
||||
- Include prerequisite Reads and post-hoc verification steps in comparisons.
|
||||
|
||||
When a task falls entirely outside Serena's design scope (e.g., reading config files, small text edits where Edit
|
||||
already sends minimal payload), classify it as "not applicable" rather than as a negative delta. A negative delta
|
||||
requires that Serena targets the task and performs worse, not that a tool designed for something else is suboptimal when
|
||||
misapplied to it.
|
||||
|
||||
---
|
||||
|
||||
## Exploration phase — tasks to actually perform
|
||||
|
||||
Work through the following. Each item exercises a specific capability under correct use; substitute an equivalent if an item isn't applicable ("no suitable candidate in this codebase" is a valid reason).
|
||||
Work through the following. Each item exercises a specific capability under correct use; substitute an equivalent if an
|
||||
item isn't applicable.
|
||||
|
||||
### Codebase understanding
|
||||
|
||||
1. Get a high-level overview of the repository structure — top-level layout, main packages, entry points.
|
||||
2. Pick one large source file (300+ lines). Get a structural overview of it. Do it with semantic overview tools and with Glob/Grep/Read. Then write out the concrete next step on each side ("after this overview, to read the body of method X I'd call _____") and compare the pair of calls, not just the overview call. Whichever tool's output most directly feeds its own follow-up wins on workflow terms.
|
||||
2. Pick one large source file (300+ lines). Get a structural overview of it. Do it with semantic overview tools and with
|
||||
Glob/Grep/Read. Then write out the concrete next step on each side and compare the pair of calls, not just the
|
||||
overview call.
|
||||
3. Pick a specific method inside a class and retrieve its body without reading the surrounding file.
|
||||
4. For one non-trivial symbol, find all references across the codebase. Compare recall and precision under the question "who uses this in code?" vs "where is this mentioned anywhere, including docs?" — these are different questions, and each toolset is naturally suited to one of them.
|
||||
5. For a class, list its subclasses / implementations and its supertypes, including transitively. Compare against what text search would need to do (and whether it can follow an override chain or cross into stub files in one step).
|
||||
6. For at least one symbol from an external dependency (a third-party library), try to retrieve its definition or signature. Note whether each toolset can do this at all and what infrastructure it requires (env activation, site-packages discovery, etc.).
|
||||
4. For one non-trivial symbol, find all references across the codebase. Compare recall and precision under the
|
||||
question "who uses this in code?" vs "where is this mentioned anywhere, including docs?"
|
||||
5. For a class, list its subclasses / implementations and its supertypes, including transitively. Compare against what
|
||||
text search would need to do.
|
||||
6. For at least one symbol from an external dependency (a third-party library), try to retrieve its definition or
|
||||
signature. Note whether each toolset can do this at all and what infrastructure it requires (environment activation,
|
||||
site-packages discovery, language-server indexing, etc.).
|
||||
|
||||
### Single-file edits — span the full range of edit sizes
|
||||
|
||||
Do all three sizes below, not just one. The comparison between content-anchored editing (Edit) and symbol-body replacement is size-dependent: Edit's payload grows with the old+new anchor pair, while symbolic body-replace grows with the full new body. They cross over, and where they cross is the whole point. Only testing small edits hides the crossover.
|
||||
7a. Small tweak (1–3 lines inside a method). Change an error message or rename a local variable inside a larger method.
|
||||
Do it with `Edit` and with symbolic body replacement. Compare payload sent, payload received, and prerequisite reads.
|
||||
|
||||
7a. Small tweak (1–3 lines inside a method). Change an error message or rename a local variable inside a larger method. Do it with `Edit` and with `replace_symbol_body`. Compare payload sent, payload received, and prerequisite reads.
|
||||
7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the main logic of a method while keeping its
|
||||
signature. Do it both ways.
|
||||
|
||||
7b. Medium rewrite (replace ~10–30 lines — most of a method body). Rewrite the main logic of a method while keeping its signature. Do it both ways. Note that for Edit you may need a long anchor to keep the old_string unique, and that for symbolic replacement the new body is similar in size to what you'd send for Edit's new_string alone.
|
||||
7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire body. Do it both ways.
|
||||
|
||||
7c. Large/whole-body rewrite. Pick a method of 50+ lines and rewrite the entire body. Do it both ways. This is where symbolic body replacement is designed to win: Edit has to send the entire old body as an anchor and the entire new body, while `replace_symbol_body` sends only the new body + a short name path. Measure the ratio.
|
||||
|
||||
8. Insert a new function/method at a specific structural location (e.g., "right after this existing method"). Try both the symbolic-insert path and the manual Edit path.
|
||||
8. Insert a new function/method at a specific structural location (for example, right after an existing method). Try
|
||||
both the symbolic-insert path and the manual Edit path.
|
||||
9. Rename a private helper used only within one file. Compare doing it by hand vs. using a semantic rename.
|
||||
|
||||
### Multi-file changes
|
||||
|
||||
10. Rename a symbol (function, class, or method) used across several files including imports. Compare the semantic path against the built-in equivalent chain. Under correct use, a semantic rename is paired with a short post-rename `Grep` for text-surface references (docstrings, markdown, notebooks) — count that as a complementary step, not a Serena failure.
|
||||
11. Move a symbol from one module to another, updating imports at all call sites. Use the semantic move tool if available; plan the built-in equivalent honestly (how many Reads, Edits, and import-cleanup decisions would it take?).
|
||||
12. Delete a symbol safely, checking it has no remaining usages. Compare "search-then-delete" with a safe-delete tool.
|
||||
13. Inline a small helper into its call sites — only if the codebase contains a function that is legally inlinable (single-expression body, no early returns, no side effects, substitutable at its call sites). If no such candidate exists, report "no suitable candidate" and skip. Do not contrive a broken input.
|
||||
10. Rename a symbol (function, class, or method) used across several files including imports. Compare the semantic path
|
||||
against the built-in equivalent chain.
|
||||
11. Move a symbol from one module to another, updating imports at all call sites. Use the semantic move tool if
|
||||
available; plan the built-in equivalent honestly.
|
||||
12. Move a file or package to a different location, updating imports at all call sites. Use the semantic move tool if
|
||||
available; plan the built-in equivalent honestly.
|
||||
12. Delete a symbol safely, checking it has no remaining usages. Compare search-then-delete with a safe-delete tool.
|
||||
13. Delete a symbol and propagate the deletion to all call sites. Compare to how the built-in equivalent would work.
|
||||
13. Inline a small helper into its call sites — only if the codebase contains a function that is legally inlinable. If
|
||||
no such candidate exists, report "no suitable candidate" and skip it.
|
||||
|
||||
### Reliability & correctness under correct use
|
||||
|
||||
14. Scope precision. Demonstrate that semantic tools address symbols by name path and can target a specific class method, override, or overload that text search would over-match. The point is to show the capability — precision under correct use — not to manufacture a rename that breaks polymorphism through careless targeting.
|
||||
15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are updated or none. A chain of `Edit` calls is not. You don't have to force a failure — report on what this means for reliability on real failures (disk full, interrupted process, transient permission errors).
|
||||
16. Success signals. For each completed refactor, note what each tool returns on success. Both toolsets report mechanical success only; semantic intent is always the caller's responsibility to verify with a diff. Note this as a baseline for both sides, not as a weakness of either.
|
||||
14. Scope precision. Demonstrate that semantic tools address symbols by name path and can target a specific class
|
||||
method, override, or overload that text search would over-match.
|
||||
15. Atomicity. A semantic cross-file refactoring is atomic: either all sites are updated or none. A chain of `Edit`
|
||||
calls is not.
|
||||
16. Success signals. For each completed refactor, note what each tool returns on success.
|
||||
|
||||
### Workflow effects across multiple edits
|
||||
|
||||
17. Chain at least three edits in one file. Report what each toolset requires between edits. Pay particular attention to whether identifiers survive mutation: name paths stay valid across edits in other regions of the file; line numbers and byte offsets don't. This is the single biggest workflow-level efficiency effect.
|
||||
18. Multi-step exploration across the repo. Note whether intermediate results (overviews, reference lists) remain useful across later edits, or whether they have to be refreshed. Stable output survives a session; ephemeral output does not.
|
||||
17. Chain at least three edits in one file. Report what each toolset requires between edits.
|
||||
18. Multi-step exploration across the repo. Note whether intermediate results remain useful across later edits or have
|
||||
to be refreshed.
|
||||
|
||||
### Things where the comparison shouldn't be interesting
|
||||
|
||||
19. Read and understand a non-code file (config, changelog, docs, notebook). Semantic-code tools don't apply — use `Read`. State the applicability boundary once and move on.
|
||||
20. Search for a free-text pattern across the repo (log string, magic constant, URL). Use `Grep`. Don't call symbolic search on free text.
|
||||
19. Read and understand a non-code file (config, changelog, docs, notebook). Semantic-code tools don't apply — use
|
||||
`Read`.
|
||||
20. Search for a free-text pattern across the repo (log string, magic constant, URL). Use `Grep`.
|
||||
|
||||
---
|
||||
|
||||
## Evaluation phase
|
||||
|
||||
Write a report structured for progressive disclosure — lead with the strongest insights; a reader should be able to stop at any point and still walk away informed.
|
||||
Write a report structured for progressive disclosure.
|
||||
|
||||
**Value-weighting is not optional.** For every contribution you name — in §1, §2, §7, and anywhere else you list what Serena adds — you must estimate *how much it matters in general coding work*, not just whether it's novel. A capability that saves 10 calls but is used once a month is a smaller practical contribution than a capability that saves 1 call but is used every editing session, and a reader needs to be able to tell which kind of contribution each item is. Be explicit about **frequency** (how often does this matter in typical Python coding?) and **value per hit** (when it matters, how much does it save?). Order by the product, not by how impressive the individual feature sounds.
|
||||
**Value-weighting is required.** For every contribution or difference you identify — positive, neutral, or negative —
|
||||
estimate:
|
||||
|
||||
**Every section must end with a one-sentence verdict** (a short paragraph labelled "**Verdict:** ...") that gives the reader the single-sentence takeaway for that section. This applies to every top-level section §1–§9 and to each subsection under §3. The verdict is a recommendation in context — e.g. "use symbolic addressing whenever you expect multiple edits to one file," "every task in this group is a clear Serena win," "built-ins only; this is not a contest." Not a hedge, not a summary — a pointed one-liner a reader can act on.
|
||||
- **Frequency:** how often this arises in typical coding work
|
||||
- **Value per hit:** calls saved, tokens saved, or correctness impact
|
||||
|
||||
1. **Headline: what Serena adds.** Open with a sharp description of the delta Serena provides on top of the built-ins — not a thumbs-up/thumbs-down, not a "they're complementary" hedge, but a specific list of what Serena contributes that built-ins alone cannot. Structure it as a short list of *capabilities* (things that become possible) and *efficiency multipliers* (things that get cheaper by how much), **ordered by how much value each contribution actually delivers in general coding work — not by novelty**. For each item, state frequency (how often does this matter?) and value per hit (when it matters, how much?). Group items into high/medium/low value tiers if the gap is large enough that a reader should care about the ordering. A reader stopping after this section should know not only what the added value is but also how much of their daily work it touches. If your opening paragraph could be written about a toolset that adds *nothing*, rewrite it. End with a one-sentence verdict.
|
||||
2. **Added value, by area (3–6 bullets).** Each bullet answers: *what would a built-ins-only workflow be missing here, and how much would it miss it?* Lead with the areas of largest weighted value — not the most novel capability, but the one whose frequency × value-per-hit product is biggest. Each bullet must include a concrete frequency estimate and a concrete value-per-hit estimate (in calls saved, tokens saved, or correctness improved). Bullets should describe what Serena contributes, not "wins" or "losses" against the built-ins. If you catch yourself writing "Serena wins at X," rewrite as "Serena adds X, which shows up in [frequency] coding work and saves [value]." End with a one-sentence verdict.
|
||||
3. **Detailed evidence, grouped by capability.** Per-task: what you tried, the full end-to-end call chain on each side (including prerequisite reads and complementary follow-up steps), payloads sent and received. Be specific — "1 call vs ~10, and the Serena call sent ~200 tokens while the Edit equivalent would have sent ~450 after including the prerequisite Read" is useful; "Serena was faster" is not. End each subsection with its own one-sentence verdict in context (e.g. "Verdict (multi-file refactors): every task in this group is a clear Serena win").
|
||||
4. **Token-efficiency analysis.** Separate from raw call count. Address:
|
||||
- Payload asymmetry as a function of edit size. Where's the crossover between content-anchored editing and symbolic body replacement? Show the ratio at small, medium, and large edit sizes.
|
||||
- Forced reads: does a tool make you load content into context that you don't actually need?
|
||||
- Stable vs ephemeral addressing. Name paths survive edits; line numbers don't. The output-size comparison must account for shelf life, not just raw tokens returned.
|
||||
Order findings by **frequency × value-per-hit**, not novelty.
|
||||
|
||||
End with a one-sentence verdict on when token economics favor each toolset.
|
||||
5. **Reliability & correctness analysis** — under correct use. Address:
|
||||
- Precision of matching: semantic identifiers vs textual matches, with concrete cases where the question shape determined which tool was the right one.
|
||||
- Scope disambiguation across override chains and overloads.
|
||||
- Atomicity of multi-file operations, and what it means on real failures.
|
||||
- Transitive semantic queries (type hierarchy, reference chains, dependency lookups) that text search cannot approximate in one step.
|
||||
**Every section must end with a one-sentence verdict** summarizing the practical takeaway.
|
||||
|
||||
End with a one-sentence verdict on where correctness weight favors each toolset.
|
||||
6. **Workflow effects across a multi-step session.** How do efficiency gaps widen over a session? Do identifiers survive edits? Does intermediate output (overviews, reference lists) stay useful? This is often where the real gap lives — a one-call comparison can miss it. End with a one-sentence verdict.
|
||||
7. **Capabilities with no built-in equivalent.** What Serena made possible that the built-ins genuinely cannot do, or can only approximate with a much longer workflow. Name each capability individually **and annotate each with how much it matters in general coding work** — a capability that's unique but rarely needed is a smaller contribution than one that's unique and needed constantly. End with a one-sentence verdict.
|
||||
8. **Where built-ins remain the right default.** Which tasks are still best served by Grep/Read/Edit, and why. Include the non-code-file boundary and the post-rename text sweep explicitly. Estimate what share of daily coding these cases represent — this is what calibrates §1's weighting. End with a one-sentence verdict.
|
||||
9. **Usage rule for a developer with both toolsets.** A per-task decision rule: given both toolsets installed, which do you reach for and when? This is the practical takeaway, not the headline — the description of added value belongs in §1. End with a one-sentence verdict.
|
||||
---
|
||||
|
||||
### 1. Headline: what Serena changes
|
||||
|
||||
Open with a precise description of the delta Serena provides.
|
||||
Distinguish between three categories:
|
||||
(a) tasks where Serena adds capability,
|
||||
(b) tasks where Serena applies but offers
|
||||
no improvement, and (c) tasks outside Serena's scope.
|
||||
Only category (b) constitutes a neutral or negative finding.
|
||||
Category (c) is context, not a finding.
|
||||
|
||||
A reader stopping here should understand both what is gained and what is not.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
### 2. Added value and differences by area (3–6 bullets)
|
||||
|
||||
Each bullet must describe:
|
||||
|
||||
- What Serena changes relative to built-ins (positive, neutral, or negative)
|
||||
- Frequency
|
||||
- Value per hit
|
||||
|
||||
Avoid framing in terms of “wins”; describe concrete differences.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
### 3. Detailed evidence, grouped by capability
|
||||
|
||||
Per task:
|
||||
|
||||
- What you attempted
|
||||
- Full call chain on both sides
|
||||
- Payloads sent and received
|
||||
|
||||
Include cases where:
|
||||
|
||||
- Serena is better
|
||||
- Built-ins are better
|
||||
- No meaningful difference exists
|
||||
|
||||
End each subsection with a verdict.
|
||||
|
||||
---
|
||||
|
||||
### 4. Token-efficiency analysis
|
||||
|
||||
Address:
|
||||
|
||||
- Payload differences across edit sizes
|
||||
- Forced reads
|
||||
- Stable vs ephemeral addressing
|
||||
|
||||
Include cases where each toolset is more efficient.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
### 5. Reliability & correctness (under correct use)
|
||||
|
||||
Address:
|
||||
|
||||
- Precision of matching
|
||||
- Scope disambiguation
|
||||
- Atomicity
|
||||
- Semantic queries vs text search
|
||||
- External dependency symbol lookup and what setup it depends on
|
||||
|
||||
Include both strengths and limitations of each toolset.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
### 6. Workflow effects across a session
|
||||
|
||||
Evaluate multi-step workflows and whether advantages compound or diminish. Include neutral or negative findings where
|
||||
applicable.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
### 7. Unique capabilities (if any)
|
||||
|
||||
List capabilities that have no practical built-in equivalent. If none exist, explicitly state that. Annotate each with
|
||||
frequency and impact.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
### 8. Tasks outside Serena's scope (built-in only)
|
||||
|
||||
Identify tasks where built-ins are the natural choice because Serena's tools don't target them. List these briefly for
|
||||
completeness but do not frame them as Serena shortcomings — they are outside its scope. Estimate their share of daily
|
||||
work to contextualize how much of a session Serena's augmentation covers.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
### 9. Practical usage rule
|
||||
|
||||
Provide a decision rule for choosing between toolsets based on task type.
|
||||
|
||||
**Verdict:** (one sentence)
|
||||
|
||||
---
|
||||
|
||||
## What I'm looking for
|
||||
|
||||
Ground every claim in something you actually measured or observed. If an initial impression turns out to be wrong as you gather more evidence, update the report.
|
||||
- Claims grounded in observed evidence
|
||||
- Explicit reporting of **positive, neutral, and negative deltas**
|
||||
- Clear quantification of impact
|
||||
- Honest workflow-level comparisons
|
||||
|
||||
**Pay particular attention to the "what does it add, and how much" question.** The §1 headline is the part a reader will actually remember. Before you write it, ask yourself: if a developer reads only that section, will they know (a) what specific capabilities and efficiency wins Serena contributes, and (b) how much those contributions actually matter in general coding work? A reader should be able to tell the difference between "rare but large" contributions and "common but small" ones without guessing. If the answer is no, rewrite it.
|
||||
|
||||
Also pay attention to:
|
||||
|
||||
- **Weighted value, not raw novelty.** A capability that saves 10 calls but is used once a month is a smaller practical contribution than one that saves 1 call but shows up every editing session. Your §1 and §2 ordering must reflect frequency × value-per-hit, not how impressive the individual feature sounds. Name both axes explicitly for each contribution.
|
||||
- **Second-order effects visible only across a session.** Token cost of re-sending content, whether identifiers survive mutation, atomicity on failure, intermediate output staying valid. A one-call comparison misses these; a multi-edit session exposes them. These are often high-frequency contributions that look small per hit.
|
||||
- **Workflow-level honesty.** Compare the full call chain on each side, not just the first call. Don't let one toolset's vocabulary set the evaluation axes.
|
||||
- **Capability deltas under correct use.** What does Serena let you do that you genuinely couldn't do — or couldn't do cheaply — with built-ins alone? These are the findings that justify the heaviest weights in §1.
|
||||
---
|
||||
|
||||
## What I am not looking for
|
||||
|
||||
- **Failures from misuse.** Calling a code-semantic tool on a non-code file, inlining an un-inlineable function, renaming without a valid name path, passing a body that drops a branch to `replace_symbol_body` without first reading the current body. These are caller errors, not tool findings.
|
||||
- **Gotcha hunts and "symmetric failure mode" comparisons.** If you find yourself writing "tool X silently does Y when misused" or "tool A has a loud failure that tool B lacks," stop. You've drifted from evaluation into user guide. A destructive write accepting what you send it is a contract, not a flaw; an addressing scheme that forces you to see what you're addressing is a consequence of addressing, not a safety feature. Report capability and efficiency, not caller-discipline requirements.
|
||||
- **Mid-session friction from tool mixing.** If a workflow only breaks because you're switching tool families mid-edit on the same file, that's a usage choice, not a tool finding.
|
||||
- **Transient glitches after `git checkout --` or other file-system mutations.** Retry once; if it then succeeds, don't include it.
|
||||
- **Hedged or neutral §1 openings.** "Both toolsets have their place" and "they're complementary, not rivals" are information-free openings that could be written about any tool pair. §1 must describe Serena's specific contribution — capabilities and efficiency multipliers — in a way that would not be true of a toolset that adds nothing. If your opening paragraph could be recycled for another tool, rewrite it.
|
||||
- **Binary adoption pitches.** "Install it" and "don't bother" are also the wrong shape of answer. The question isn't whether to adopt — it's what the added value actually is, and how much of it. A good §1 lets the reader decide for themselves whether the described delta is worth it for their work, by being specific about what the delta is and how often it shows up.
|
||||
- **Novelty-weighted ordering.** Listing Serena's most exotic capability first because it's the most impressive one misleads a reader about what they'll actually experience. If the most exotic feature is used once a month and a boring one is used every session, the boring one comes first. A contribution's weight is frequency × value-per-hit, not how surprising it is.
|
||||
- **Unquantified "wins."** "Serena is faster here," "Serena is more reliable," "built-ins are cheaper" are all unhelpful without a quantity attached. Every claim about relative value needs a concrete number or a concrete frequency — calls saved, tokens saved, times per session, share of daily edits. An evaluation without magnitudes is a review, not a measurement.
|
||||
- Misuse-based failure analysis
|
||||
- Gotcha comparisons
|
||||
- Neutral statements without evidence
|
||||
- Binary recommendations
|
||||
- Novelty-driven ordering
|
||||
- Unquantified claims
|
||||
@@ -3,7 +3,7 @@
|
||||
_Use the prompt below to summarize the evaluation of Serena by your agent.
|
||||
We used it to create the recommendations we display in the readme._
|
||||
|
||||
Read the evaluation reperot and write a one-sentence user-facing summary about the value that Serena's tools provide for coding. With a bit
|
||||
Write a one-sentence user-facing summary about the value that Serena's tools provide for coding. With a bit
|
||||
of emotion but grounded in the evaluation that you have done above. It should be written from the perspective of a coding AI agent and make it
|
||||
clear whether an AI agent would ask its owner to add Serena's tools or not. It will be the first sentence that a potential new user will see
|
||||
about how useful Serena is for an AI agent.
|
||||
about how useful Serena is for an AI agent.
|
||||
@@ -1,339 +1,354 @@
|
||||
# Evaluation by Opus 4.6 in Claude Code
|
||||
# Evaluation Report: Serena's Tools vs Built-In Tools
|
||||
|
||||
## Final Recommendation
|
||||
> **One-line summary:** Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up.
|
||||
|
||||
The recommendation is created with the [follow-up summary prompt](../011_followup-summary-prompt), everything else was created in one shot from the
|
||||
[evaluation prompt](../010_evaluation-prompt).
|
||||
|
||||
> "If I could ask my owner for one upgrade to my coding toolkit, it would be Serena — its symbolic addressing eliminated the constant re-reading
|
||||
that quietly eats my context on multi-edit sessions, its atomic cross-file refactorings collapsed 10-call manual workflows into one, and after a
|
||||
rigorous hands-on evaluation I can say that while built-in tools remain the right choice for small edits and text searches, Serena fills the exact
|
||||
gaps that make me feel clumsy without it."
|
||||
|
||||
|
||||
Ground-truth hands-on evaluation run on `tianshou` (Python, ~26 kLOC under `tianshou/`), focused on what Serena's tools add on top of `Read/Edit/Write/Glob/Grep` under correct use. Every measurement below comes from an actual tool call on this repo, not a thought experiment; edits were applied and reverted via `git checkout --` between experiments.
|
||||
**Evaluated by:** Claude Opus 4.6 (coding AI agent in Claude Code CLI)
|
||||
**Codebase:** [tianshou](https://github.com/thu-ml/tianshou) — a Python reinforcement learning library (~26K lines, 43 source files)
|
||||
**Method:** Hands-on, side-by-side execution of 20 tasks using both toolsets. All edits were applied to real files and verified via `git diff`, then reverted.
|
||||
|
||||
---
|
||||
|
||||
## §1. Headline — what Serena adds
|
||||
## 1. Headline: What Serena Changes
|
||||
|
||||
On a typical Python editing session, Serena contributes four things built-ins cannot reproduce cheaply or at all. Ranked by **frequency × value-per-hit**, highest first:
|
||||
Serena adds an IDE-backed semantic layer (powered by JetBrains) on top of the built-in text-level tools. The practical delta breaks down into three categories:
|
||||
|
||||
**Tier A — touches almost every editing session**
|
||||
**(a) Tasks where Serena adds capability:**
|
||||
- **Cross-file refactoring** (rename, move symbol, move file): 1 atomic call vs. N×(Grep+Read+Edit) chains. This is Serena's strongest contribution — it collapses multi-file, import-aware operations into single calls that are atomic and semantically correct.
|
||||
- **Structural code navigation**: Symbol overviews, type hierarchies, and reference finding return structured, scope-disambiguated results that text search cannot produce without manual interpretation.
|
||||
- **External dependency introspection**: Serena can look up symbols in installed packages via the IDE index without requiring manual environment discovery.
|
||||
|
||||
1. **Stable name-path addressing that survives intermediate edits.** Every Serena editing and querying tool is addressed by `ClassName/method_name`, which remains valid after inserts, deletes, and rewrites elsewhere in the same file. Grep/Read/Edit chain through line numbers or text anchors, both of which drift. **Frequency: every multi-edit session, which is most of them.** **Value per hit: saves one Read + one re-grep per subsequent edit in the same file (~400–800 bytes of re-sent content each time).** This is the quietest contribution and probably the largest in aggregate.
|
||||
**(b) Tasks where Serena applies but offers no improvement:**
|
||||
- **Single-file rename of a unique string**: `Edit` with `replace_all=true` achieves the same result in 1 call (after a Read), with comparable effort.
|
||||
- **Small edits (1–3 lines inside a method)**: Edit's substring matching is more efficient — less payload sent, no requirement to supply the full symbol body.
|
||||
- **Inserting new code at a known location**: Both approaches require ~1–2 calls. Edit can target by surrounding text; Serena targets by name path. The effort is comparable.
|
||||
|
||||
2. **Payload asymmetry on medium-and-larger body rewrites.** `replace_symbol_body` sends `name_path + new_body`, while `Edit` sends `old_body + new_body`. Measured crossover is around 10–15 lines: below that, Edit's tiny anchor wins 5–10× on payload; above that, Serena wins ~2× at 20 lines and ~2.3× at 66 lines, and the gap keeps widening with body size. **Frequency: once per session you rewrite a non-trivial method body.** **Value per hit: 50% payload reduction on medium rewrites, climbing toward ~60% on large ones, plus no need for a prior Read to capture the exact anchor.**
|
||||
**(c) Tasks outside Serena's scope (built-in only):**
|
||||
- Non-code files (config, docs, notebooks, changelogs)
|
||||
- Free-text search across the repo
|
||||
- Shell commands, git operations, test execution
|
||||
- File creation from scratch
|
||||
|
||||
**Tier B — rare per session but very high value when it fires**
|
||||
|
||||
3. **Single-call cross-file refactorings with atomicity and automatic import cleanup.** `rename`, `move`, and `safe_delete` each collapse an 8–12-call manual workflow (grep → read-all-callers → edit each → verify → hand-clean unused imports) into one call, atomically. **Frequency: 0–3 times per session, depending on task type — zero on many days, constant on refactoring days.** **Value per hit: 5–10× call reduction plus correctness improvements that manual chains drop (I observed `move` auto-remove a `Categorical` import from the source file that a human refactorer would easily miss).**
|
||||
|
||||
4. **Semantic navigation into third-party dependency source.** `find_symbol(search_deps=True)` / `find_declaration` resolves a type from its use site straight into `site-packages` or stub files, returning the full class body in one call. Built-ins would require: (a) reading the file to find the import, (b) locating the venv, (c) guessing the module path, (d) reading the file. **Frequency: a few times per session when debugging type errors or reviewing unfamiliar library APIs.** **Value per hit: 3–4 manual calls and a path-guessing step collapsed into one.**
|
||||
|
||||
**Tier C — capability-level, not efficiency-level**
|
||||
|
||||
5. **Transitive type hierarchy and reference graphs.** `type_hierarchy` returns the full super/sub chain — including into external stub files (`.pyi`) — in one call. `find_referencing_symbols` returns each reference *with its containing symbol* (the function/class that holds it), which Grep cannot produce. Built-ins can imitate single-level relationships with Grep but cannot chase override chains in one step. **Frequency: once or twice a session on unfamiliar code; close to never once you know the hierarchy.** **Value per hit: one call vs. N iterative Greps, plus recall into stub files.**
|
||||
|
||||
**Verdict:** Serena's biggest practical contribution is the *quietest* one — stable symbolic addressing that survives in-file mutations — and its most *spectacular* one — single-call atomic cross-file refactorings — is rare but decisive when it fires; the middle-weight wins are payload asymmetry on body rewrites and cheap dependency navigation.
|
||||
**Verdict:** Serena's primary contribution is collapsing multi-file, semantically-aware operations (rename, move, reference-finding, type hierarchy) from multi-step manual processes into single atomic calls; for single-file text edits, it is comparable or slightly less efficient.
|
||||
|
||||
---
|
||||
|
||||
## §2. Added value, by area
|
||||
## 2. Added Value and Differences by Area
|
||||
|
||||
Ordered by frequency × value-per-hit, not novelty.
|
||||
**1. Cross-file rename/move (Serena: strong positive)**
|
||||
- Serena: 1 call updates definition + all imports + all usages across N files atomically.
|
||||
- Built-in: 1 Grep + N×(Read+Edit) = 2N+1 calls, non-atomic.
|
||||
- Frequency: Moderate (a few times per working session during refactoring work). Value per hit: High — saves 6–10+ calls and eliminates the risk of partial updates.
|
||||
|
||||
- **Stable identifiers across chained edits to one file.** Demonstrated in Task 17: three sequential edits to `CollectStats` (`insert_after refresh_len_stats`, `insert_after refresh_return_stats`, `replace_symbol_body refresh_std_array_stats`) all used the same kind of stable name-path address and worked without re-reading the file. An equivalent `Edit` chain after the first two inserts would have needed at least one intermediate `Read` because line numbers and some text anchors shift. **Frequency: every multi-edit session.** **Value: ~1 re-read (~400–800 bytes) saved per subsequent edit; compounds across a session.**
|
||||
**2. Structural overview / targeted symbol reading (Serena: moderate positive)**
|
||||
- `get_symbols_overview(depth=1)`: 1 call returns all classes, methods, and attributes in structured JSON, via the IDE's parser — language-agnostic and always correct. Built-in equivalent: `Grep` for a language-specific heuristic pattern like `^(class |def )`. This works for simple Python files but is inherently fragile: the pattern must be hand-tuned per language, misses decorated or multi-line declarations, and can match inside strings or comments. For non-Python languages, an entirely different regex would be needed.
|
||||
- `find_symbol(include_body=True)`: Retrieves a specific method body by name path without reading surrounding code. Built-in: requires knowing the line number (via Grep) then Read with offset — 2 calls.
|
||||
- Frequency: High (many times per session). Value per hit: Low-to-moderate — saves ~1 call and some context window tokens per navigation. The reliability advantage over heuristic Grep patterns is consistent across all languages.
|
||||
|
||||
- **Body replacement at medium/large sizes.** Measured in Tasks 7a/7b/7c on `collector.py`:
|
||||
- 1-line change (7a): Edit `~120 B`, `replace_symbol_body` `~1200 B`. **Edit wins ~10×.**
|
||||
- ~20-line rewrite (7b): Edit `~1430 B`, Serena `~715 B`. **Serena wins ~2×.**
|
||||
- ~66-line rewrite (7c): Edit `~4700 B`, Serena `~2050 B`. **Serena wins ~2.3×.**
|
||||
Crossover sits around 10–15 lines. **Frequency: once or twice per substantive editing session.** **Value: 50–60% payload cut on the edits where it applies; grows with body size.**
|
||||
**3. Reference finding with structural context (Serena: moderate positive)**
|
||||
- `find_referencing_symbols`: Returns which *symbols* reference a target, with file grouping and usage type (import, call, type annotation). Built-in `Grep` returns all text matches including docs, comments, and string literals — same recall but lower precision.
|
||||
- Frequency: Moderate. Value per hit: Moderate — precision matters when planning refactors of widely-used symbols.
|
||||
|
||||
- **Atomic cross-file operations.** Task 10 renamed `EpisodeRolloutHookMCReturn` across `collector.py` + `test_collector.py` in one `rename` call, and — importantly — *also* updated a Sphinx `:class:` cross-reference in a docstring that a pure `Edit` chain would only catch if the caller remembered to sweep docstrings separately. Task 11 moved `get_stddev_from_dist` from `collector.py` to `batch.py`: one call updated the target file, added the import at the source, updated a separate test file's import, *and* removed the now-unused `Categorical` import from the source. **Frequency: 0–3 times per session.** **Value per hit: 5–10 calls saved plus one or two automatic correctness wins the manual path tends to miss.**
|
||||
**4. Type hierarchy (Serena: positive, niche)**
|
||||
- 1 call returns full super/sub-type chains transitively. Built-in: requires iterative Grep for `class X(Y)` patterns, manual transitive closure.
|
||||
- Frequency: Low (occasional during architecture exploration). Value per hit: High per occurrence — several calls saved.
|
||||
|
||||
- **Dependency navigation.** Task 6 resolved `Distribution` from its use site in `collector.py` into the real `torch.distributions.distribution.Distribution` class body (~300 lines) with a single `find_declaration` call. The built-in path requires parsing imports, guessing the venv site-packages location, and reading the file by hand. **Frequency: a few times per session on unfamiliar code or type debugging.** **Value: 3–4 calls plus one guess collapsed to one.**
|
||||
**5. Stable addressing across edits (Serena: moderate positive)**
|
||||
- Serena addresses symbols by name path (`CollectStats/refresh_return_stats`), which is invariant under edits elsewhere in the file. The built-in workflow is fundamentally line-number-mediated: Grep returns line numbers, Read takes line offsets — and both go stale after any edit to the file. Edit's text matching (`old_string`) is more resilient than line numbers, but it sits at the end of a chain that starts with position-based lookups. After an edit shifts lines, previously noted Grep results and Read offsets are invalid and must be re-acquired.
|
||||
- This means a multi-edit session with built-ins requires re-Grep or re-Read between edits to re-establish positions, while Serena's name paths remain valid throughout. In Task 17 (three chained edits), Serena needed 3 calls total; the built-in path needed 4 (1 Read + 3 Edits) only because the text anchors happened to be unique — but had any Edit failed uniqueness, a re-Read would have been required, pushing the count to 5–7.
|
||||
- Frequency: High (any session with multiple edits to the same file). Value per hit: Moderate — saves 1–2 re-read calls per edit chain, and eliminates the class of errors where stale line numbers cause edits to land in the wrong place.
|
||||
|
||||
- **Reference and hierarchy queries that distinguish "code usage" from "any text match".** Task 4 on `Collector`: `find_referencing_symbols` returned ~70 files with each reference tagged by containing symbol (`Py:IMPORT_ELEMENT`, `Py:FUNCTION_DECLARATION: ["test_dqn"]`), while `Grep \bCollector\b` returned 332 matches across 88 files including notebooks, SVGs, `CHANGELOG.md`, and `README.md`. These are answers to *different questions*, and each tool is naturally matched to one. **Frequency: several times per session; the right question determines the right tool.** **Value: precision saves a read of each false-positive file (~10–30 saved reads on a popular class).**
|
||||
**6. Single-file small/medium edits (Serena: neutral to slight negative)**
|
||||
- For a 1-line change inside a 13-line method, Serena's `replace_symbol_body` requires sending the entire 13-line body. Edit sends just the changed line. The prerequisite cost differs: Serena needs `find_symbol(include_body=True)` to get the body; Edit needs `Read` of the relevant lines. Both are 1 prerequisite call + 1 edit call.
|
||||
- Frequency: Very high. Value per hit: Slightly negative for Serena — more payload for small changes.
|
||||
|
||||
- **Safe-delete with automatic usage check.** Task 12: `safe_delete` on `EpisodeRolloutHookMerged` succeeded silently (no usages); on `CollectStats` it refused with a 200+ line usage list grouped by enclosing symbol. The built-in equivalent is a `Grep` followed by manual discipline — same call count but no enforcement. **Frequency: rare (delete-by-name operations).** **Value: when it fires, it prevents the most expensive class of mistake (deleting a referenced symbol).**
|
||||
|
||||
**Verdict:** The two highest-weighted contributions are the boring one (stable in-file addressing, every session) and the refactoring one (rare but 5–10× call reduction); Serena's more exotic capabilities (hierarchy, dependency navigation) are lower-frequency polish on top.
|
||||
**Verdict:** Serena's strongest contributions are in cross-file refactoring (high value, moderate frequency) and structured navigation (moderate value, high frequency); for small text edits, built-ins are slightly more efficient.
|
||||
|
||||
---
|
||||
|
||||
## §3. Detailed evidence, grouped by capability
|
||||
## 3. Detailed Evidence, Grouped by Capability
|
||||
|
||||
### §3.1 File structural overview
|
||||
### 3.1 Structural Overview (Task 2)
|
||||
|
||||
**Task 2** — structural overview of `tianshou/data/collector.py` (1551 lines).
|
||||
- Serena `get_symbols_overview(depth=0)`: ~240 bytes returned; top-level classes + functions names only, no line numbers, no per-class method detail.
|
||||
- Serena `get_symbols_overview(depth=1)`: ~3200 bytes; classes with nested method and field lists.
|
||||
- Grep `^(class |def | def )`: ~3300 bytes; identical structural information plus line numbers.
|
||||
- Output size is roughly equivalent at `depth=1`. The real difference shows up in the *follow-up* call:
|
||||
- Serena follow-up (`find_symbol("Collector/_compute_action_policy_hidden", include_body=True)`): 1 call, returns body only, ~66 lines. Address `Collector/_compute_action_policy_hidden` remains stable across any edits to unrelated regions.
|
||||
- Built-in follow-up (`Read offset=707 limit=66`): 1 call, returns body + line-number prefixes. Address (line 707) goes stale after any insert above that line.
|
||||
- The winner depends on what comes next: if you stop at reading, they tie; if you plan to edit the body next, Serena's address is the one that survives the subsequent mutation.
|
||||
| Metric | Serena `get_symbols_overview(depth=1)` | Built-in `Grep` for class/def |
|
||||
|---|---|---|
|
||||
| Calls | 1 | 1 |
|
||||
| Output structure | JSON tree: 14 classes with nested methods + attributes | Flat list: 54 `class`/`def` lines with line numbers |
|
||||
| Nesting visible? | Yes — methods grouped under classes | No — indentation implies nesting but no grouping |
|
||||
| Attributes visible? | Yes | No |
|
||||
| Output size | ~2.5KB structured JSON | ~3KB flat text |
|
||||
| Correctness | Always correct — uses IDE parser, language-agnostic | Heuristic — the regex `^(class \|def \| def )` is Python-specific and can miss decorated methods, multi-line signatures, or match inside strings/comments |
|
||||
| Language portability | Works unchanged for any language the IDE supports | Requires a new hand-tuned regex per language |
|
||||
| Next step | `find_symbol("ClassName/method", include_body=True)` — 1 call | `Read(file, offset=line, limit=N)` — 1 call |
|
||||
|
||||
**Verdict (§3.1):** tie on the overview call alone; Serena wins on the full read-then-edit chain because of address stability.
|
||||
Both paths need 1 call for overview + 1 call for drill-down. Serena's output is more structured and reliably correct; Grep's is simpler but fragile. The Grep pattern used here happened to work well for this Python codebase, but would need to be rewritten for Java (`class|interface|enum`, brace-delimited), Rust (`fn|struct|impl|trait`), TypeScript (`class|function|interface`), etc. — and even within Python, it misses cases like `@overload`-decorated methods or methods whose `def` line is preceded by a long decorator stack that pushes `def` to a non-standard indentation.
|
||||
|
||||
### §3.2 Method body retrieval
|
||||
**Verdict:** Serena provides a reliably correct structural overview across languages; Grep-based overviews are heuristic approximations that work in simple cases but degrade for complex declarations or non-Python codebases.
|
||||
|
||||
**Task 3** — body of `Collector._compute_action_policy_hidden`.
|
||||
- Serena: 1 call (`find_symbol include_body`), ~1800 bytes of body-only output.
|
||||
- Built-ins: 1 call (`Read offset=707 limit=66`), ~1900 bytes including line prefixes.
|
||||
- Essentially equivalent on a cold cache; Serena's output is cleaner for machine consumption, built-ins' is cleaner for humans reviewing line locations.
|
||||
### 3.2 Targeted Method Retrieval (Task 3)
|
||||
|
||||
**Verdict (§3.2):** tie in isolation; Serena wins when followed by an edit.
|
||||
| Metric | Serena | Built-in |
|
||||
|---|---|---|
|
||||
| Calls | 1 (`find_symbol` with `include_body=True`) | 2 (Grep to find line + Read with offset) |
|
||||
| Prerequisite | Know the name path | Know the method name |
|
||||
| Output | Method body only (no surrounding code) | Surrounding code included (must estimate limit) |
|
||||
|
||||
### §3.3 Reference search
|
||||
For `Collector/_collect` (330 lines): Serena returned exactly 330 lines. Built-in Read returned 332 lines (including the next method's signature).
|
||||
|
||||
**Task 4** — who uses `Collector`?
|
||||
- Serena `find_referencing_symbols`: ~70 files, each reference annotated with its containing symbol. Precise to code usage. Missed docstring/README mentions almost entirely (except a few `FILE`-tagged entries).
|
||||
- Grep `\bCollector\b`: 332 occurrences across 88 files, includes `docs/*.ipynb`, `docs/*.md`, `structure.svg`, `CHANGELOG.md`, `README.md`.
|
||||
- Cost comparison for the question "every code file that *uses* this class":
|
||||
- Serena: 1 call, answer usable directly.
|
||||
- Grep: 1 call, answer needs filtering by extension and visual dedup.
|
||||
- Cost comparison for "anywhere this name appears in the repo, including docs":
|
||||
- Serena: cannot answer directly.
|
||||
- Grep: 1 call, answer usable directly.
|
||||
**Verdict:** Serena saves 1 call and returns precisely the requested body; built-in is adequate but requires line-number discovery.
|
||||
|
||||
**Verdict (§3.3):** each toolset wins on its own question; pick by question shape, and do not penalize either for missing the other's answer.
|
||||
### 3.3 Cross-File References (Task 4)
|
||||
|
||||
### §3.4 Type hierarchy and override chains
|
||||
| Metric | Serena `find_referencing_symbols` | Built-in `Grep` |
|
||||
|---|---|---|
|
||||
| Calls | 1 | 1 |
|
||||
| Files found | 63 (code files only, structured by symbol type) | 83 (includes docs, notebooks, changelogs, README) |
|
||||
| Output structure | Grouped by file, categorized (import, call, type annotation) | Flat file list |
|
||||
| Precision | Code-only references | All textual mentions |
|
||||
|
||||
**Task 5** — super/sub types of `BaseCollector`.
|
||||
- Serena `type_hierarchy(depth=0)`: 1 call, returns transitive hierarchy including `Collector → AsyncCollector` on the sub side and `ABC → object` on the super side — the super side is resolved into an external `<ext:abc.pyi>` stub file.
|
||||
- Grep `class \w+\(.*BaseCollector`: 1 call, returns exactly one hit (`Collector`). To find `AsyncCollector` I'd have to issue a second grep for `class \w+\(.*Collector`, and so on recursively. External supertypes cannot be reached at all.
|
||||
- **This is a capability delta, not a raw efficiency delta**: text search cannot cross an override chain in one step, and cannot reach `.pyi` stubs.
|
||||
Grep finds 83 files including `README.md`, `CHANGELOG.md`, and `.ipynb` notebooks. Serena finds 63 code files. For the question "who uses this in code?", Serena's result is directly usable. For "where is this mentioned anywhere?", Grep is the right tool.
|
||||
|
||||
**Verdict (§3.4):** clear Serena capability win; value grows with hierarchy depth and cross-file reach.
|
||||
**Verdict:** Serena provides higher-precision code-usage results; Grep provides broader text-level coverage. Different tools for different questions.
|
||||
|
||||
### §3.5 Dependency navigation
|
||||
### 3.4 Type Hierarchy (Task 5)
|
||||
|
||||
**Task 6** — resolve `Distribution` from its usage in `collector.py`.
|
||||
- Serena `find_declaration` with a regex anchored at the usage site: 1 call, returned the full 300-line body of `torch/distributions/distribution.py::Distribution`, including docstrings. Unambiguous — the tool used the import context to pick the right file out of 41 candidates named `Distribution` in the dependency tree.
|
||||
- Built-in equivalent: (a) Read `collector.py` imports, (b) map `from torch.distributions import Distribution` to `torch/distributions/distribution.py`, (c) locate the venv's site-packages path, (d) Read the file. 3–4 calls plus one implicit "where is the venv" step.
|
||||
| Metric | Serena `type_hierarchy` | Built-in |
|
||||
|---|---|---|
|
||||
| Calls | 1 | 2+ (Grep for `class X(BaseCollector` + Grep for `class Collector(` + potential transitive search) |
|
||||
| Result | `BaseCollector → ABC → object` (supers), `BaseCollector → Collector → AsyncCollector` (subs), with file locations | `class Collector(BaseCollector[TCollectStats], ...)` at line 551 — requires further reads for AsyncCollector's superclass |
|
||||
|
||||
**Verdict (§3.5):** clear Serena capability win whenever you need third-party source.
|
||||
**Verdict:** Serena produces the full transitive hierarchy in 1 call; built-in requires iterative searches.
|
||||
|
||||
### §3.6 Small edits (< ~10 lines)
|
||||
### 3.5 External Dependency Lookup (Task 6)
|
||||
|
||||
**Task 7a** — one-line error message change inside `BaseCollector._validate_buffer` (21-line body).
|
||||
- `Edit old_string=".. should be greater than 0." new_string=".. must be strictly positive."`: ~120 bytes on the wire, one call. Prerequisite: know the exact anchor (one prior Grep or Read, which I already had from earlier in the session).
|
||||
- `replace_symbol_body`: ~1200 bytes (whole 21-line body resent). Prerequisite: prior `find_symbol include_body` to see the current body.
|
||||
- Result diff identical in both cases (`1 insertion(+), 1 deletion(-)`).
|
||||
| Metric | Serena | Built-in |
|
||||
|---|---|---|
|
||||
| Can retrieve? | Yes — `find_declaration` + `search_deps=True` | Requires environment discovery: `python -c "import X; print(inspect.getfile(X))"` then `Read` |
|
||||
| Infrastructure | IDE index (pre-built) | Working Python environment, correct venv activated |
|
||||
| Result | Symbol location + documentation | Full source file (if env is set up) |
|
||||
|
||||
**Verdict (§3.6):** Edit wins ~10× on payload for ≤ 3-line changes; use Edit whenever the change is small and the anchor is obvious.
|
||||
In this session, the Python environment wasn't directly accessible from bash, so built-in lookup would have required additional setup. Serena retrieved `torch.distributions.Distribution`'s location and docstring via the IDE index.
|
||||
|
||||
### §3.7 Medium edits (~10–30 lines)
|
||||
**Verdict:** Serena provides dependency introspection without environment setup; built-in requires a working interpreter.
|
||||
|
||||
**Task 7b** — rewrite ~20 lines inside `CollectStats.update_at_step_batch`.
|
||||
- Edit: ~750 B old anchor + ~680 B new body = ~1430 B on the wire.
|
||||
- `replace_symbol_body`: ~680 B new body + ~35 B name_path = ~715 B.
|
||||
- Both end with identical diffs.
|
||||
- Prerequisite is symmetric: both need to see the current body first (one `find_symbol` or one `Read`).
|
||||
### 3.6 Small Edit — 1 Line Change (Task 7a)
|
||||
|
||||
**Verdict (§3.7):** crossover point — Serena pulls ahead around 15 lines and wins ~2× at 20.
|
||||
| Metric | Serena `replace_symbol_body` | Built-in `Edit` |
|
||||
|---|---|---|
|
||||
| Prerequisite calls | 1 (`find_symbol` with `include_body=True`) — already done in Task 3 | 1 (`Read` of ~15 lines) |
|
||||
| Edit call | 1 (send full 13-line body) | 1 (send 1-line old + 1-line new) |
|
||||
| Payload sent (edit call) | ~550 chars (full body) | ~120 chars (changed line only) |
|
||||
| Total calls | 1–2 | 2 |
|
||||
|
||||
### §3.8 Large edits (50+ lines)
|
||||
**Verdict:** For small changes, Edit sends ~4.5× less payload in the edit call. Total call count is the same.
|
||||
|
||||
**Task 7c** — whole-body rewrite of `Collector._compute_action_policy_hidden` (~66 lines).
|
||||
- Edit: ~2700 B old anchor + ~2000 B new body = ~4700 B.
|
||||
- `replace_symbol_body`: ~2000 B new body + name_path = ~2050 B.
|
||||
- Ratio ~2.3×. The asymmetry grows linearly with body size: Edit scales as `O(old + new)`, `replace_symbol_body` as `O(new)`.
|
||||
### 3.7 Medium Rewrite — ~19 Lines (Task 7b)
|
||||
|
||||
**Verdict (§3.8):** clear Serena win; reach for `replace_symbol_body` on any substantial method-body rewrite.
|
||||
| Metric | Serena | Built-in |
|
||||
|---|---|---|
|
||||
| Prerequisite | 1 `find_symbol` (already done) | 1 `Read` (~22 lines) |
|
||||
| Edit payload | ~550 chars (full body) | ~1000 chars (old 19 lines + new 20 lines) |
|
||||
| Total calls | 1–2 | 2 |
|
||||
|
||||
### §3.9 Structural insertion
|
||||
At this scale, Serena's payload is actually *smaller* because Edit must send both old and new text, while Serena sends only the new body. The crossover point is approximately where the changed region exceeds half the method body.
|
||||
|
||||
**Task 8** — insert a new method right after `CollectStats.refresh_len_stats`.
|
||||
- Serena `insert_after_symbol` with `name_path="CollectStats/refresh_len_stats"` and body of the new method: ~150 B on the wire, unambiguous anchor, 1 call. Result: new method inserted cleanly between `refresh_len_stats` and `refresh_std_array_stats`.
|
||||
- Edit equivalent: old_string must capture the end of `refresh_len_stats` uniquely — roughly the last 5 lines of the method — then new_string replays those lines and appends the new method. Approximately 550 B on the wire.
|
||||
**Verdict:** At medium scale, payloads converge; Serena's is slightly smaller due to only sending the replacement.
|
||||
|
||||
**Verdict (§3.9):** Serena wins ~3× on payload and on anchor stability for structural inserts.
|
||||
### 3.8 Large Rewrite — 55 Lines (Task 7c)
|
||||
|
||||
### §3.10 Single-file rename of a private helper
|
||||
| Metric | Serena | Built-in |
|
||||
|---|---|---|
|
||||
| Prerequisite | 1 `find_symbol` (already done) | 1 `Read` (~67 lines) |
|
||||
| Edit payload | ~2200 chars (new body) | ~4400 chars (old 63 lines + new 62 lines) |
|
||||
| Total calls | 1–2 | 2 |
|
||||
|
||||
**Task 9** — rename `_HACKY_create_info_batch` → `_create_info_batch_legacy` (2 occurrences, same file).
|
||||
- Serena `rename` via name_path: 1 call.
|
||||
- Edit with `replace_all=true`: 1 call. Both the declaration and the one call site get rewritten.
|
||||
- Both succeed; both leave a clean 2-line diff.
|
||||
**Verdict:** For full-method rewrites, Serena sends ~50% less payload because it only sends the replacement, not the original.
|
||||
|
||||
**Verdict (§3.10):** tie — for single-file renames of a distinctive identifier, built-in `Edit(replace_all=true)` is competitive.
|
||||
### 3.9 Cross-File Rename (Task 10)
|
||||
|
||||
### §3.11 Multi-file rename
|
||||
| Metric | Serena `rename` | Built-in chain |
|
||||
|---|---|---|
|
||||
| Calls | 1 | 1 Grep + 4 Read + 4 Edit = 9 |
|
||||
| Files affected | 4 (automatically discovered) | 4 (manually discovered via Grep) |
|
||||
| Import handling | Automatic | Manual — must identify and rewrite each import statement |
|
||||
| Atomicity | Atomic — all-or-nothing | Sequential — intermediate states are inconsistent |
|
||||
|
||||
**Task 10** — rename `EpisodeRolloutHookMCReturn` → `EpisodeRolloutMCReturnHook` (5 sites across `tianshou/data/collector.py` and `test/base/test_collector.py`).
|
||||
- Serena `rename`: 1 call, atomic across both files. Also updated a Sphinx `:class:` docstring cross-reference at `collector.py:611` via the default `rename_in_comments=True`.
|
||||
- Built-in chain: `Grep` (1) → `Edit replace_all=true` on `collector.py` (1) → `Edit replace_all=true` on `test_collector.py` (1) → verification `Grep` (1). 4 calls minimum, 5–6 if I remember to hit the docstring reference. Not atomic across files.
|
||||
**Verdict:** Serena reduces a 9-call chain to 1 call for cross-file rename, with atomicity.
|
||||
|
||||
**Verdict (§3.11):** Serena wins ~4–5× on call count and catches the docstring reference by default.
|
||||
### 3.10 Move Symbol Across Modules (Task 11)
|
||||
|
||||
### §3.12 Cross-module move
|
||||
| Metric | Serena `move` | Built-in equivalent |
|
||||
|---|---|---|
|
||||
| Calls | 1 | ~8–12 (Read source, Edit source to remove, Edit target to add, Grep for imports, Read+Edit each import site) |
|
||||
| Import updates | Automatic | Manual — must rewrite `from X import Y` → `from Z import Y` in each file |
|
||||
| Dependency imports | Automatically adds needed imports to target file | Must manually inspect what the moved function imports and replicate |
|
||||
|
||||
**Task 11** — move `get_stddev_from_dist` from `tianshou/data/collector.py` to `tianshou/data/batch.py`.
|
||||
- Serena `move`: 1 call. Final `git diff --stat`:
|
||||
```
|
||||
test/base/test_stats.py | 3 ++-
|
||||
tianshou/data/batch.py | 24 ++++++++++++++++++++++++
|
||||
tianshou/data/collector.py | 27 ++-------------------------
|
||||
```
|
||||
The tool:
|
||||
1. Inserted the function into `batch.py`.
|
||||
2. Removed it from `collector.py`.
|
||||
3. Added `from tianshou.data.batch import get_stddev_from_dist` to `collector.py`.
|
||||
4. Updated `test/base/test_stats.py`'s import of this function.
|
||||
5. **Removed the now-unused `Categorical` import from `collector.py`**, because nothing else in that file referenced it.
|
||||
- Built-in chain (honestly planned): Grep for callers (1) → Read each caller to see current import form (≥ 3) → Edit collector.py to remove the function (1) → Edit batch.py to insert (1) → Edit each caller's import (≥ 2) → Edit collector.py to remove the now-unused `Categorical` import (1) → verification Grep (1). **Roughly 10 calls**, and the "unused import" cleanup is the kind of thing a distracted human misses.
|
||||
Moving `get_stddev_from_dist` from `collector.py` to `stats.py`: Serena updated 3 files (source, target, test) in 1 call, including adding necessary imports to the target module.
|
||||
|
||||
**Verdict (§3.12):** biggest call-count collapse in the evaluation — ~10× — with a correctness bonus.
|
||||
**Verdict:** Move is Serena's highest-value single operation — it handles import graph updates that would be error-prone manually.
|
||||
|
||||
### §3.13 Safe deletion
|
||||
### 3.11 Move File (Task 12a)
|
||||
|
||||
**Task 12** — two deletions.
|
||||
- `safe_delete(EpisodeRolloutHookMerged)` (unused): succeeded, removed the class cleanly, 38-line deletion.
|
||||
- `safe_delete(CollectStats)` (heavily used): refused with an explicit `SafeDeleteFailedException`, returning ~220 usage locations grouped by enclosing symbol.
|
||||
- Built-in equivalent: `Grep` for usages → visual inspection → `Edit` deletion. Same call count when the symbol is unused; similar when it isn't, but with no enforcement — a careless `Edit` would happily delete a used symbol and break the repo.
|
||||
| Metric | Serena `move` | Built-in equivalent |
|
||||
|---|---|---|
|
||||
| Calls | 1 | 1 `git mv` + 1 Grep + N×(Read+Edit) for import updates = ~12 |
|
||||
| Files updated | 5 import updates automatic | Must manually discover and rewrite |
|
||||
|
||||
**Verdict (§3.13):** same call count as manual, but the enforcement eliminates the worst mistake class.
|
||||
**Verdict:** Same pattern as symbol move — Serena collapses an N-step process.
|
||||
|
||||
### §3.14 Inline helper
|
||||
### 3.12 Safe Delete (Task 12b)
|
||||
|
||||
**Task 13** — skipped. No legally inlinable helper (single-expression, side-effect-free, with call sites) found in a quick scan of `tianshou/`. Per the prompt, I did not contrive a broken input.
|
||||
Serena's `safe_delete` in safe mode (default):
|
||||
- Reports all usages before deleting — acts as a guard.
|
||||
- For unused symbols, deletes cleanly in 1 call.
|
||||
- For used symbols, refuses and lists usages.
|
||||
|
||||
**Verdict (§3.14):** no data; comparison not applicable on this codebase.
|
||||
Built-in equivalent: Grep to check usages → if none found, Read + Edit to delete. 2–3 calls.
|
||||
|
||||
### §3.15 Scope precision and disambiguation
|
||||
**Verdict:** Safe delete adds a safety check that the built-in path must implement manually.
|
||||
|
||||
**Task 14** — find every `_collect` in `collector.py`.
|
||||
- `find_symbol("_collect")` returned three distinct hits with their full name paths (`BaseCollector/_collect`, `Collector/_collect`, `AsyncCollector/_collect`), plus inlined signatures for each. Each is addressable unambiguously — I can rename, replace-body, or reference exactly one of them.
|
||||
- `Grep "_collect\("` would return all three locations but with no structural distinction: to tell them apart I'd have to read surrounding context to see which class each `def` belongs to.
|
||||
### 3.13 Inline (Task 13)
|
||||
|
||||
**Verdict (§3.15):** Serena's addressing is precise by construction on override chains and overloads.
|
||||
Serena's inline tool did not work for any Python function tested. The JetBrains backend does not support Python function inlining (this is a language-level limitation of the IDE's refactoring engine, not a Serena bug).
|
||||
|
||||
### §3.16 Chained edits to one file
|
||||
**Verdict:** No candidate inlinable in this codebase; inline capability appears unavailable for Python.
|
||||
|
||||
**Task 17** — three successive edits on `CollectStats`:
|
||||
1. `insert_after_symbol("CollectStats/refresh_len_stats", ...)` — insert new `reset_len_stats` method.
|
||||
2. `insert_after_symbol("CollectStats/refresh_return_stats", ...)` — insert new `reset_return_stats` method.
|
||||
3. `replace_symbol_body("CollectStats/refresh_std_array_stats", ...)` — rewrite an existing method body.
|
||||
### 3.14 Scope Precision (Task 14)
|
||||
|
||||
All three calls used the original name paths, unchanged; none required a re-Read between edits. Final `git diff` showed exactly the expected three-edit composite: two insertions plus one body rewrite, adjacent and clean. The first two inserts shifted line numbers between 8 and 16 lines, which would have invalidated any line-number addressing for the third target.
|
||||
Three methods named `reset_env` exist in `collector.py` (in `BaseCollector`, `Collector`, `AsyncCollector`).
|
||||
|
||||
An equivalent Edit chain would have survived this particular sequence (because the anchors were text, not line numbers), but it exposes the *general* pattern: name_paths are mutation-proof addresses, line numbers are not, and large text anchors become non-unique quickly.
|
||||
- Serena: `find_symbol("Collector/reset_env")` returns exactly that override's body.
|
||||
- Grep: Returns 3 line numbers. Determining which belongs to which class requires reading surrounding context.
|
||||
|
||||
**Verdict (§3.16):** for any session with three or more edits to the same file, symbolic addressing is a systematic safety and efficiency win.
|
||||
**Verdict:** Serena's name-path addressing eliminates class-level disambiguation that Grep requires.
|
||||
|
||||
### §3.17 Non-code files and free-text searches
|
||||
### 3.15 Chained Edits (Task 17)
|
||||
|
||||
**Tasks 19/20** — state the applicability boundary and move on. Semantic tools don't apply to changelogs, notebooks, configs, or free-text searches for log strings; `Read` and `Grep` are the right tools.
|
||||
Three consecutive edits to methods in `CollectStats`:
|
||||
- Serena: 3 `replace_symbol_body` calls, 0 intermediate reads. Name paths (`CollectStats/refresh_return_stats`, `CollectStats/refresh_len_stats`, `CollectStats/refresh_std_array_stats`) remained valid across all edits because they are structural identifiers, not positions.
|
||||
- Built-in: 1 Read + 3 Edit calls = 4 calls in the best case. The Read was required upfront (Edit enforces "must read before editing"). The 3 Edits succeeded without re-reads only because the `old_string` anchors happened to remain unique after each edit. But this is fragile: Edit also enforces a "file has been modified since read" check when external tools modify the file, forcing a re-Read. More fundamentally, if I had needed to *find* these methods first (the typical case when you don't already know the line numbers), each Grep result from before the first edit would have been stale after it — line 232 is no longer line 232 after inserting 2 lines above it.
|
||||
|
||||
**Verdict (§3.17):** built-ins only; not a contest.
|
||||
The core asymmetry: Serena's addressing is *structural* (survives edits by definition), while built-in addressing is *positional* (line numbers from Grep/Read go stale after any insertion or deletion). Edit's text matching partially mitigates this, but only for the final step — the discovery steps (Grep, Read with offset) remain position-dependent.
|
||||
|
||||
**Verdict:** Serena's name-path stability eliminates the re-read/re-grep cycle between edits; built-in tools require re-acquiring positions after each edit that shifts line numbers.
|
||||
|
||||
---
|
||||
|
||||
## §4. Token-efficiency analysis
|
||||
## 4. Token-Efficiency Analysis
|
||||
|
||||
**Payload asymmetry by edit size** (measured on `collector.py`):
|
||||
### Payload by edit size
|
||||
|
||||
| Edit size | Edit (old+new) | `replace_symbol_body` | Ratio |
|
||||
| Edit scale | Serena payload (edit call) | Edit payload (edit call) | Winner |
|
||||
|---|---|---|---|
|
||||
| 1 line (in 21-line method) | ~120 B | ~1200 B | **Edit 10× smaller** |
|
||||
| ~20 lines rewritten | ~1430 B | ~715 B | **Serena 2× smaller** |
|
||||
| ~66 lines whole-body | ~4700 B | ~2050 B | **Serena 2.3× smaller** |
|
||||
| 1-line change in 13-line method | ~550 chars (full body) | ~120 chars | Edit (~4.5×) |
|
||||
| 19-line medium rewrite | ~550 chars | ~1000 chars | Serena (~1.8×) |
|
||||
| 55-line full rewrite | ~2200 chars | ~4400 chars | Serena (~2×) |
|
||||
| Cross-file rename (4 files) | ~100 chars | ~800 chars (4 Edit calls) | Serena (~8×) |
|
||||
|
||||
Crossover is around 10–15 lines. Below it, Edit's per-change payload is dominated by the tiny anchor and wins by an order of magnitude. Above it, Edit pays once for the old body and again for the new, while `replace_symbol_body` pays only for the new body plus a ~30-character name_path; the gap grows linearly with body size. Structural inserts (`insert_after_symbol`) have similar asymmetry — the name_path replaces a multi-line text anchor.
|
||||
### Prerequisite reads
|
||||
|
||||
**Forced reads.** Serena's overview tools return symbol names without bodies and its reference tools return containing-symbol metadata without snippets, so you control when to pull code into context. `find_symbol(include_body=False)` + `find_symbol(include_body=True, name_path=...)` is a two-step "browse, then fetch body" pattern that keeps context lean; the built-in equivalent is `Grep` (which does not return bodies) + `Read` with an offset/limit, which is about as lean but requires the caller to compute the limit by hand.
|
||||
- Serena: `find_symbol(include_body=True)` returns the symbol body. If you've already navigated to it during exploration, no additional call needed.
|
||||
- Built-in: `Read` with offset/limit. Always required before Edit (enforced by the tool).
|
||||
|
||||
**Stable vs ephemeral addressing.** Name paths remain valid across edits to unrelated regions of the same file. Line numbers and byte offsets do not, and text anchors become ambiguous once a file grows. The output-size comparison has to account for *shelf life*: a slightly larger overview that stays useful across an entire session is cheaper than a slightly smaller one that has to be regenerated after each edit. In a five-edit session on one file, Serena's name-path overview is queried once; the line-number grep output is effectively refreshed after each edit that shifts upstream content — an O(edits × file_grep_cost) hidden tax that the one-call comparison misses.
|
||||
### Stable vs ephemeral addressing
|
||||
|
||||
**Verdict (§4):** under ~10-line edits, built-in `Edit` is cheaper; above that threshold, symbolic body replacement wins on raw payload; across a multi-edit session, name-path addressing wins regardless of size because it doesn't decay.
|
||||
- Serena's name paths (`CollectStats/refresh_return_stats`) are stable identifiers — they survive any edit to surrounding code. No re-read or re-discovery needed between edits.
|
||||
- The built-in workflow uses ephemeral, position-based addresses at every stage: Grep returns line numbers, Read takes line offsets, and both are invalidated by any insertion or deletion above the target. Edit's `old_string` matching is content-based and more resilient, but it depends on the upstream position-based steps to know *what* to match. After an edit shifts line numbers, the entire Grep→Read→Edit chain must be re-executed from the top.
|
||||
- In a session with N edits to the same file, this costs up to N-1 additional Grep/Read round-trips with built-ins. With Serena, the cost is zero — the same name path works on the first and tenth edit.
|
||||
|
||||
**Verdict:** Serena is more token-efficient for medium-to-large edits and cross-file operations; Edit is more efficient for small, localized changes. Across multi-edit sessions, Serena's stable addressing avoids the re-read tax that compounds with each successive built-in edit.
|
||||
|
||||
---
|
||||
|
||||
## §5. Reliability and correctness analysis (under correct use)
|
||||
## 5. Reliability & Correctness (Under Correct Use)
|
||||
|
||||
**Precision of matching.** `find_referencing_symbols` on `Collector` returns ~70 code files and annotates each with the containing symbol that holds the reference. `Grep \bCollector\b` returns 332 hits across 88 files including notebooks, SVGs, and markdown. For the question "which Python files import and use this class?", Serena's output is directly usable and Grep's needs a filter pass. For the question "where is the name `Collector` mentioned anywhere in the repo, including docs, changelog, and diagrams?", Grep's output is directly usable and Serena's is incomplete by design. Each tool's precision is perfect *for its question*; the mistake is asking the wrong one.
|
||||
### Precision of matching
|
||||
- Serena: Exact symbol resolution via name paths. `Collector/reset_env` unambiguously selects one method.
|
||||
- Edit: Text matching. Unique strings match correctly; non-unique strings fail (and Edit reports the error).
|
||||
|
||||
**Scope disambiguation across overrides.** `find_symbol("_collect")` returned three distinct name paths — `BaseCollector/_collect`, `Collector/_collect`, `AsyncCollector/_collect` — each independently addressable for rename or body replacement. Text search on `_collect(` returns three line locations with no structural annotation; the caller must read context to tell them apart, and any cross-file rename risks touching the wrong override if called carelessly.
|
||||
### Scope disambiguation
|
||||
- Serena distinguishes overrides, overloads (via indices), and nested classes by name path.
|
||||
- Grep/Edit cannot distinguish methods with the same name in different classes without reading surrounding context.
|
||||
|
||||
**Atomicity on real failures.** `move` of `get_stddev_from_dist` made coordinated changes across three files in one call. A five-step Edit chain replicating the same move would leave the repo in a half-renamed state if any intermediate step failed on disk-full, a permission error, or an interrupted process; a single-call atomic refactoring either completes or leaves the working tree clean. This matters less for local agent sessions (where the blast radius of a partial refactor is small and recoverable with `git checkout --`) and more for any workflow where a failed run is committed or pushed.
|
||||
### Atomicity
|
||||
- Serena's cross-file operations (rename, move) are atomic — all files updated or none.
|
||||
- Built-in multi-file edits are sequential — a failure mid-chain leaves an inconsistent state (recoverable via `git checkout`).
|
||||
|
||||
**Transitive semantic queries.** `type_hierarchy` returned both the sub-chain (`Collector → AsyncCollector`) and the super-chain (`BaseCollector → ABC → object`, with `ABC` resolved into an external `.pyi` stub) in one call. No text-search workflow reaches into stub files or chains override relationships in one step.
|
||||
### Semantic queries vs text search
|
||||
- `find_referencing_symbols` returns code-level references categorized by type. Grep returns all textual mentions.
|
||||
- `type_hierarchy` returns transitive sub/supertypes. No built-in equivalent without iterative search.
|
||||
|
||||
**Success signals (symmetric).** Both toolsets return only mechanical success: "the file was written," "the rename finished." Neither verifies that the new code still compiles, type-checks, or preserves semantics. Post-edit `git diff` review is the caller's responsibility on both sides.
|
||||
### External dependency lookup
|
||||
- Serena: Available via IDE index, no environment setup needed. Can retrieve symbol docs and location.
|
||||
- Built-in: Requires working Python environment, correct venv, and manual file discovery. More powerful when available (full source), but higher setup cost.
|
||||
|
||||
**Verdict (§5):** Serena's correctness edge is concentrated in questions that are semantic by nature — override chains, transitive type queries, cross-file atomic ops — and tied to Grep/Edit on questions that are textual by nature.
|
||||
**Verdict:** Serena provides stronger correctness guarantees for symbol-level operations (scope, atomicity, semantic precision); built-in tools are reliable for text-level operations with the caveat that scope disambiguation requires manual effort.
|
||||
|
||||
---
|
||||
|
||||
## §6. Workflow effects across a multi-step session
|
||||
## 6. Workflow Effects Across a Session
|
||||
|
||||
The single biggest session-level effect is **identifier stability across edits**. A session that makes 5 edits to `collector.py` looks like:
|
||||
### Compound advantages
|
||||
- **Exploration → Edit without position re-acquisition**: Serena's overview tools produce name paths that are directly usable as edit targets. The built-in path produces line numbers (from Grep) that are consumed by Read — but after an edit, those line numbers are stale. In a typical explore-edit-explore-edit cycle, Serena's name paths remain valid throughout while built-in line numbers must be re-acquired after each edit. Over 3 explore+edit cycles, this saves ~3–6 intermediate Grep/Read calls.
|
||||
- **Chained edits without re-reads**: Name-path stability means no re-reads between edits to the same file. For N edits, this saves up to N-1 Read calls. Edit's text matching partially avoids this (if old_strings stay unique), but the upstream discovery steps (Grep line numbers, Read offsets) still go stale.
|
||||
- **Cross-file refactoring**: When a rename or move is part of a larger change, doing it atomically avoids the need to manually track "which files still need updating."
|
||||
|
||||
- With symbolic addressing: one `get_symbols_overview` at the start, five `replace_symbol_body`/`insert_after_symbol` calls by name path. The overview is consulted zero or one more times. No re-reads of the file between edits.
|
||||
- With line-number or large-text-anchor addressing: one `Grep class/def` at the start, one `Read offset+limit` before each edit (or a careful re-grep after any insert that shifts lines), five `Edit`s. The Grep/overview may need to be refreshed mid-session once line numbers drift.
|
||||
### Diminishing returns
|
||||
- For purely exploratory sessions (reading code, no edits), Serena's advantage is modest — `Grep` and `Read` are nearly as fast for navigation, and Serena's structured output doesn't save many calls.
|
||||
- For sessions dominated by small text edits (config changes, log message tweaks), Serena adds no value.
|
||||
|
||||
The hidden cost of the built-in workflow is not in any single call — it's the compounding re-Reads and anchor recomputation across a session. A single Read of a 1500-line file is ~40 kB of context; doing it four extra times across a session is a ~160 kB invisible tax that never shows up in a one-call comparison.
|
||||
### Neutral findings
|
||||
- Both toolsets require similar total calls for single-file work. The difference is ~1 call per operation.
|
||||
- Output quality for code review (reading diffs, understanding changes) is identical — both require `git diff`.
|
||||
|
||||
**Intermediate output survives.** The overview I pulled at the start of this evaluation, the reference list for `Collector`, the type hierarchy for `BaseCollector`, and the signature table for `_collect` all remain valid now that I'm writing the report — I never had to regenerate them. A workflow based on line numbers would have had to regenerate its intermediate output after each of the ~15 edits I applied and reverted during the experiments.
|
||||
|
||||
**Verdict (§6):** session-level efficiency scales with mutation rate; the more edits you plan, the more decisively symbolic addressing wins, and the effect is invisible on any single-call benchmark.
|
||||
**Verdict:** Serena's advantages compound across multi-step refactoring sessions where each operation feeds into the next; for read-heavy or small-edit sessions, the advantage is marginal.
|
||||
|
||||
---
|
||||
|
||||
## §7. Capabilities with no built-in equivalent
|
||||
## 7. Unique Capabilities (No Practical Built-In Equivalent)
|
||||
|
||||
1. **Cross-file atomic refactorings with automatic import maintenance.** `move` cleaned up an unused import in the source file as a side effect of moving the last user of that import. Built-ins have no equivalent — you'd have to notice. **Value: rare but high — this is the kind of cleanup that bit-rots across a large refactor.**
|
||||
1. **Atomic cross-file rename** — 1 call, all imports and usages updated, all-or-nothing. No built-in equivalent without scripting a multi-step chain. Frequency: moderate (refactoring sessions). Impact: high (saves 5–10 calls, eliminates partial-update risk).
|
||||
|
||||
2. **Resolution of third-party symbols from a usage site.** `find_declaration` + `find_symbol(search_deps=True)` reach into `site-packages` and `.pyi` stubs for the exact class used at a given line of code. **Value: a few times per session, saves 3–4 calls each.**
|
||||
2. **Atomic cross-file move** (symbol or file) with import rewriting — includes adding necessary imports to the target module. Frequency: low-to-moderate. Impact: very high per occurrence (saves 8–12 calls, handles import dependency graph).
|
||||
|
||||
3. **Transitive type hierarchy including external supertypes.** `type_hierarchy` returns sub- and super-chains in one call and crosses module boundaries and stub files. No text-search sequence can reproduce this in O(1). **Value: once or twice per unfamiliar codebase; near zero once you know it.**
|
||||
3. **Type hierarchy traversal** — transitive supertypes and subtypes in 1 call. Built-in requires iterative Grep with manual transitive closure. Frequency: low. Impact: moderate (saves 3–5 calls).
|
||||
|
||||
4. **Containing-symbol metadata on reference queries.** `find_referencing_symbols` returns each reference with its enclosing function/class name, making the result a navigation map rather than a line list. **Value: every time you need to understand *how* a symbol is used, not just *where*.**
|
||||
4. **Safe delete with usage check** — reports all usages before deleting, with option to propagate. Built-in requires Grep + manual verification. Frequency: low. Impact: moderate (safety check is the value).
|
||||
|
||||
5. **Enforced safe-delete with usage-list refusal.** `safe_delete` refuses to remove a symbol that still has usages and returns the usage list. Built-ins cannot refuse — `Edit` applies whatever you send it. **Value: rare but prevents the worst-class mistake.**
|
||||
5. **External dependency symbol lookup** via IDE index — no Python environment needed. Frequency: moderate. Impact: moderate (avoids environment setup friction).
|
||||
|
||||
**Verdict (§7):** the capability deltas are real but concentrated in lower-frequency tasks; the one that shows up across *every* editing session is (1.5) — symbolic addressing as a property of the *editing* tools themselves, which I'm treating as the Tier-A efficiency win in §1 rather than a separate capability here.
|
||||
**Verdict:** Serena provides 3–5 capabilities with no practical built-in equivalent, concentrated in cross-file refactoring and semantic navigation.
|
||||
|
||||
---
|
||||
|
||||
## §8. Where built-ins remain the right default
|
||||
## 8. Tasks Outside Serena's Scope (Built-In Only)
|
||||
|
||||
- **Small anchored edits (≤ ~10 lines).** Task 7a: Edit's payload is ~10× smaller than `replace_symbol_body` for a one-line change because it sends only the two anchors, not the whole body. Frequency: extremely high — typo fixes, constant changes, error-message tweaks, single-line bug fixes. **Probably 30–50% of daily edits.**
|
||||
- **Non-code file operations**: Reading/editing config files, docs, changelogs, notebooks → `Read`/`Edit`/`Write`
|
||||
- **Free-text search**: Finding log strings, URLs, magic constants → `Grep`
|
||||
- **Shell operations**: Running tests, builds, git commands, package management → `Bash`
|
||||
- **File creation**: New files from scratch → `Write`
|
||||
- **Glob-based file discovery**: Finding files by pattern → `Glob`
|
||||
- **Broad codebase search**: When you don't know what you're looking for → `Grep` with regex
|
||||
|
||||
- **Free-text search for strings, log messages, magic constants, URLs.** Task 20: `Grep` is the only tool that can find a bare string across the repo. Serena's symbolic search expects an identifier, not a phrase. **Frequency: multiple times per session.**
|
||||
Estimated share of daily work: These tasks constitute roughly 40–60% of a typical coding session (reading docs, running tests, searching for patterns, editing config). Serena's augmentation covers the remaining 40–60% where code-level semantic operations apply.
|
||||
|
||||
- **Non-code files** — changelogs, READMEs, YAML configs, notebooks. Task 19: `Read` is the tool. **Frequency: occasional but universal.**
|
||||
|
||||
- **Doc and docstring sweeps after a code-level rename.** Serena's `rename_in_comments=True` catches Sphinx cross-references (verified in Task 10), but if your documentation lives outside of Python docstrings — `.md`, `.rst`, `.ipynb` — a text-based `Grep` sweep is the complementary step, not a Serena failure. **Frequency: every cross-file rename that touches a public API.**
|
||||
|
||||
- **Single-file renames of distinctive identifiers.** Task 9: `Edit replace_all=true` is effectively tied with semantic rename; either works. **Frequency: common.**
|
||||
|
||||
- **Quick one-shot explorations where you don't plan to edit the target.** If the workflow is "look at one function, answer a question, move on," `Read offset/limit` and `Grep` are as fast as semantic tools and don't require any address to be stable beyond the current call.
|
||||
|
||||
These cases are not rare — collectively they probably cover more than half of the calls in a typical session, which is exactly why §1's weighting puts the quiet "addressing stability" win above the spectacular "cross-file refactoring" wins: the former touches every session, the latter touches a handful per week.
|
||||
|
||||
**Verdict (§8):** built-ins are the right default for small edits, free-text search, non-code files, and docstring sweeps — roughly half of daily editing work.
|
||||
**Verdict:** Built-in tools handle roughly half of daily work that falls entirely outside Serena's scope; Serena augments the code-centric other half.
|
||||
|
||||
---
|
||||
|
||||
## §9. Usage rule for a developer with both toolsets
|
||||
## 9. Practical Usage Rule
|
||||
|
||||
Per-task decision rule, in priority order:
|
||||
| Task type | Use |
|
||||
|---|---|
|
||||
| Cross-file rename, move, delete | Serena (unique capability) |
|
||||
| Understanding class hierarchy or symbol relationships | Serena (1 call vs. iterative search) |
|
||||
| Getting a structural overview of a large file | Serena (richer structure) or Grep (simpler, faster) |
|
||||
| Reading a specific method body by name | Serena (direct) or Grep+Read (2 calls) |
|
||||
| Small edit (1–5 lines inside a method) | Edit (less payload) |
|
||||
| Full method rewrite | Serena `replace_symbol_body` (less payload, no re-read) |
|
||||
| Single-file rename of a unique identifier | Edit with `replace_all` (equivalent) |
|
||||
| Non-code files, config, docs | Built-in `Read`/`Edit` |
|
||||
| Text search, pattern matching | Built-in `Grep` |
|
||||
| Anything involving shell, git, tests | Built-in `Bash` |
|
||||
| External dependency inspection | Serena (no env setup needed) |
|
||||
|
||||
1. **Small edit (≤ ~10 lines), known text anchor** → `Edit`. Payload is ~10× smaller than symbolic body replacement at this size.
|
||||
2. **Medium or larger body rewrite, or structural insert** → `replace_symbol_body` / `insert_after_symbol`. 2–3× payload cut plus stable addressing for the next edit.
|
||||
3. **Cross-file rename, move, or delete of a symbol** → `rename` / `move` / `safe_delete`, then a complementary `Grep` sweep of `.md`/`.rst`/`.ipynb` for any text-only references. The semantic tool covers code and Python docstrings in one atomic step; the Grep sweep handles external docs.
|
||||
4. **Find callers of a class/function** → `find_referencing_symbols`.
|
||||
5. **Find any mention of a name across the repo including docs** → `Grep`.
|
||||
6. **Navigate into third-party library source** → `find_declaration` with a regex anchored at the use site, or `find_symbol(search_deps=True)`.
|
||||
7. **Understand a type hierarchy** → `type_hierarchy`.
|
||||
8. **Understand an unfamiliar file's structure** → `get_symbols_overview(depth=1)` for classes + methods, then `find_symbol(include_body=True)` for the methods you actually want to read.
|
||||
9. **Chained edits to one file** → address every edit by `name_path`; do not mix in line-number or large-text-anchor edits unless you plan to re-Read before each one.
|
||||
10. **Free-text search, non-code files, small single-line tweaks** → built-ins.
|
||||
|
||||
**Verdict (§9):** Edit for small changes, symbolic body replacement for larger ones, semantic multi-file refactoring tools with a Grep doc-sweep on top, Grep/Read for text and non-code; the one habit that matters most is defaulting to name-path addressing whenever you expect more than one edit to the same file.
|
||||
**Verdict:** Use Serena for cross-file refactoring and semantic navigation; use built-ins for text-level edits, non-code files, and shell operations; for single-file code edits, choose based on edit size (Edit for small, Serena for full-body rewrites).
|
||||
@@ -1,136 +1,660 @@
|
||||
# Evaluation by Gpt 5.4 in Codex CLI
|
||||
# Evaluation: What Serena Adds Over Built-Ins
|
||||
|
||||
## Final Recommendation
|
||||
**Evaluated by:** Gpt 5.4 (high) in Codex
|
||||
|
||||
The recommendation is created with the [follow-up summary prompt](../011_followup-summary-prompt), everything else was created in one shot from the
|
||||
[evaluation prompt](../010_evaluation-prompt).
|
||||
This report compares Serena's JetBrains-backed semantic coding tools with built-in file, shell, search, and patch tools in this repository. The comparison assumes competent use of both toolsets: built-ins are used for text, file, shell, config, and small patch work; Serena is used where symbol identity, language semantics, or IDE refactoring semantics apply.
|
||||
|
||||
> "As a coding AI agent, I would ask my owner to add Serena because it lets me work with code as stable symbols instead of fragile text, turning the
|
||||
refactors, lookups, and multi-file edits that usually feel risky and repetitive into precise IDE-backed operations I can trust."
|
||||
Method: I explored code first, avoided repo documentation and prior notes, ran real edits/refactors, and reverted after each experiment. After each edit/refactor experiment I checked `git status --short` and returned the tree to clean before moving on. Measurements are approximate, but call counts, diff sizes, and result shapes are from observed runs.
|
||||
|
||||
## 1. Headline
|
||||
> **One-line summary:** As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and
|
||||
refactoring, which makes real code changes feel faster, safer, and far less blind.
|
||||
|
||||
Serena's largest added value is stable semantic addressing and IDE-backed refactoring. In this Java plugin repo, that meant a method/class could be targeted as `Logger/warning[0]`, `Symbol/safeDelete`, or `ProjectUtil` without line numbers, and cross-file rename/move/inline/delete operations were delegated to IntelliJ's model.
|
||||
## 1. Headline: what Serena changes
|
||||
|
||||
High value, common: symbol lookup, method-body retrieval, stable name paths, and reference search. Frequency: many times per coding session. Value per hit: usually saves 1-3 reads/searches and avoids loading whole files.
|
||||
Serena adds a semantic layer over the codebase. Its concrete delta is the ability to address and transform code by symbols: name paths, overload indexes, reference graphs, type hierarchies, external declarations, and JetBrains refactoring operations.
|
||||
|
||||
High value, less frequent but large: cross-file rename, move, safe delete, inline. Frequency: a few times per feature/refactor. Value per hit: saves roughly 5-20 calls and reduces missed import/call-site risk.
|
||||
Tasks where Serena adds capability:
|
||||
|
||||
Medium value: type hierarchy and external dependency declaration lookup. Frequency: occasional. Value per hit: turns "search and infer" into 1 semantic query; text search cannot truly reproduce transitive hierarchy or dependency source lookup without IDE/index/cache work.
|
||||
- Code-reference search: `find_referencing_symbols` for `Symbol/findReferences` returned 6 precise code usages grouped by containing symbols. `rg findReferences` also returned the declaration, the `/findReferences` route string, and comments.
|
||||
- Type hierarchy: one `type_hierarchy` call for `PostRequestHandler` returned 18 direct endpoint subclasses, a transitive `TypeHierarchyHandler -> GetSupertypesHandler/GetSubtypesHandler` branch, and the external `Object` supertype.
|
||||
- External dependency lookup: `find_declaration` resolved `ReferencesSearch.search(anchorElement)` to `<ext:ReferencesSearch.class|466808a0>`, and `find_symbol` returned the selected overload body. Built-in repo/cache grep found the import and call but not the dependency definition.
|
||||
- Cross-file refactors: semantic rename, symbol move, file move, safe delete, and inline executed real IDE refactors and updated files/usages/imports.
|
||||
- Stable addressing: a chained edit used `SymbolFinder/findFilesByName`, `SymbolFinder/getProject`, and `SymbolFinder/qualNameMatchesName` after earlier edits shifted line numbers.
|
||||
|
||||
Low/no added value: config/docs/free-text search and tiny line edits. Frequency: common, but built-ins are already optimal. Value per hit for Serena: none or negative for small local edits.
|
||||
Tasks where Serena applies but offered little or no improvement:
|
||||
|
||||
**Verdict:** Serena adds the most value whenever the task is about named code entities rather than text spans; the weighted daily win is stable symbol navigation, while the biggest per-hit win is IDE refactoring.
|
||||
- Tiny intra-method edits: a 1-line error-message change was smaller with a built-in patch; Serena body replacement required sending the whole method.
|
||||
- Simple one-file private rename: manual patch and semantic rename produced the same 2-line diff for `qualNameMatchesName -> qualifiedNameMatchesName`.
|
||||
- Method insertion at a known spot: both workflows produced the same 4-line diff; Serena's benefit was target stability, not smaller payload.
|
||||
- Whole-method rewrite: Serena targets the method boundary cleanly, but the full replacement body still has to be sent.
|
||||
|
||||
## 2. Added Value By Area
|
||||
Tasks outside Serena's scope:
|
||||
|
||||
- Stable symbol navigation: showed on `Symbol.java` and `UIControlUtil.java`. Frequency: every non-trivial coding session. Value: saves 1-3 calls per lookup and avoids full-file reads.
|
||||
- Symbol-scoped edits: `replace_symbol_body`, `insert_after_symbol`, and overload-specific targeting worked on `Logger/warning[0]`, `Logger/warning[1]`, and `Logger/logToToolWindow`. Frequency: several times per session. Value: small edits lose to text Edit, medium edits break even, whole-body edits save about 2x input payload.
|
||||
- Cross-file refactors: renaming `Logger` to `SerenaLogger` updated 10 files plus the Java file rename; moving `ProjectUtil` into `service.endpoint` updated the package and removed the now-local import. Frequency: occasional. Value: saves roughly 10-20 manual reads/edits/verifications.
|
||||
- Semantic relationships: `ToolWindowContent` references, `TypeHierarchy` subtypes, `SubtypeHierarchy` supertypes, and `Gson.fromJson` declaration came back as code entities. Frequency: occasional. Value: 1 call versus several searches plus inference.
|
||||
- Built-in text work remains essential: `build.gradle.kts`, `/findSymbol`, `127.0.0.1`, and `FORM_INIT_DELAY_MILLIS` were naturally handled by `Read`/`rg`. Frequency: large share of daily work. Value: Serena adds nothing there.
|
||||
- Reading config/non-code files such as `plugin.xml` and `build.gradle.kts`.
|
||||
- Free-text searches for URLs, endpoint strings, settings values, comments, and resources.
|
||||
- File inventory, line counts, shell commands, build/test execution, Git operations, and arbitrary text edits.
|
||||
|
||||
**Verdict:** Serena's contribution is not blanket speed; it removes repeated code-entity bookkeeping from ordinary navigation and almost all bookkeeping from real refactors.
|
||||
Value-weighted summary:
|
||||
|
||||
## 3. Detailed Evidence
|
||||
| Rank | Difference | Frequency | Value per hit | Delta |
|
||||
|---|---:|---:|---:|---|
|
||||
| 1 | Cross-file semantic refactors | Medium-high | High | Often avoids 5-20 manual edit/search/verify steps and reduces partial-update risk. |
|
||||
| 2 | Code usages vs text mentions | High | Medium-high | Removes noisy grep triage for "who uses this symbol?" |
|
||||
| 3 | Symbol overview/body retrieval | High | Medium | Avoids large reads and stale line targeting. |
|
||||
| 4 | Type hierarchy/implementations | Medium | High | Hard to reproduce correctly with text search, especially transitively. |
|
||||
| 5 | Stable addressing across edit chains | Medium | Medium | Reduces refresh work after line shifts. |
|
||||
| 6 | External dependency lookup | Low-medium | High | Needs IDE index; built-ins need source/decompiler infrastructure. |
|
||||
| 7 | Small local edits | High | Neutral/negative for Serena | Built-in patch/edit usually sends less. |
|
||||
|
||||
### 3.1 Code Understanding
|
||||
**Verdict:** Serena materially changes symbol-centric exploration and refactoring; it does not replace built-ins for ordinary text, file, shell, config, or tiny local edits.
|
||||
|
||||
Semantic overview of `Symbol.java` returned a class tree with fields, methods, inner classes, and overload indexes in 1 call. Text equivalent was `rg "class |...\\(" Symbol.java`, which returned 100+ signature/comment hits and needed filtering. Follow-up semantic read of `UIControlUtil/findButton` was 1 call returning only the 10-line body; text follow-up needed locating the line then reading a slice.
|
||||
## 2. Added value and differences by area
|
||||
|
||||
Payloads: semantic overview sent path/depth only and returned about 900 tokens; text grep returned about 2,000+ tokens for `Symbol.java`. Semantic method read returned about 90 tokens; text slice returned similar body tokens but required a locator step.
|
||||
- Cross-file semantic refactoring changes both workflow and correctness. Frequency: medium-high. Value per hit: high. The observed class rename updated a file/class plus imports/usages; the nested-record move changed 4 existing files and created 1 new file; the file move updated package/imports across 3 paths. Built-ins can reproduce the result, but only through search, file moves, patches, import cleanup, and verification.
|
||||
|
||||
**Verdict:** Use Serena for source structure and specific method bodies; use text only when the question is literally textual.
|
||||
- Semantic search separates code usage from text mention. Frequency: high. Value per hit: medium-high. Serena returned only code references for `Symbol/findReferences`; `rg` returned code, route strings, comments, and the declaration. Built-ins remain better for "mentioned anywhere."
|
||||
|
||||
### 3.2 References, Hierarchy, Dependencies
|
||||
- Symbol overview/body retrieval cuts exploration payload. Frequency: high. Value per hit: medium. `get_symbols_overview` on `Symbol.java` returned fields, methods, nested classes, and overload indexes without reading the 1,000+ line file. Regex produced line-oriented candidates and still required manual scope interpretation.
|
||||
|
||||
`find_referencing_symbols` on `ToolWindowContent` returned 4 code uses: two subclass declarations and two parameters. `rg "ToolWindowContent"` returned 5 lines including the definition. For "who uses this in code," Serena was higher precision; for "where is this string mentioned," `rg` was the right tool.
|
||||
- Type/dependency queries add capabilities not practically present in built-ins. Frequency: low-medium. Value per hit: high. Serena returned a transitive type graph and an external IntelliJ API method body; text tools needed iterative searches or extra source/decompiler setup.
|
||||
|
||||
`type_hierarchy` on `TypeHierarchy` returned `SubtypeHierarchy` and `SupertypeHierarchy`; supertypes for `SubtypeHierarchy` returned `TypeHierarchy` and external `Object`. Text search found name matches plus false positives like `TypeHierarchyRequest`.
|
||||
- Small localized edits are not improved by symbol-body replacement. Frequency: high. Value per hit: neutral or negative for Serena. A one-line patch was smaller than replacing an 11-line method body.
|
||||
|
||||
`find_declaration` on `gson.fromJson(requestBody, requestClass)` resolved external `Gson/fromJson[0]` and returned the source body. Built-ins needed finding the Gradle dependency, locating `~/.gradle/.../gson-2.10.1-sources.jar`, listing/extracting `Gson.java`, then searching inside it.
|
||||
- Some JetBrains refactors staged added/deleted files. Frequency: limited to class/file move/delete style operations. Value per hit: neutral operational tradeoff. Semantic value remains, but verification/revert must check the index as well as unstaged files.
|
||||
|
||||
**Verdict:** Serena adds real semantic reach for "code relationships"; text search can find mentions, but not reliably answer relationship questions in one step.
|
||||
**Verdict:** Serena's highest-value differences are symbol identity, reference graphs, and IDE refactor execution; built-ins remain superior for simple text-local work.
|
||||
|
||||
### 3.3 Edit Size Economics
|
||||
## 3. Detailed evidence, grouped by capability
|
||||
|
||||
Small edit: rewriting `Logger.logToToolWindow` as an early return. Text patch sent a small old/new anchor, about 9 changed lines. Serena required sending the whole 11-line method body. Text was cheaper.
|
||||
### 3.1 Repository structure and entry points
|
||||
|
||||
Medium edit: rewriting most of `UIControlUtil.tryHandleDialogs`. Text patch sent about 14 old lines plus 20 new lines. Serena sent the full method body, including its comment, about 26 lines. Roughly even.
|
||||
Attempted: identify top-level layout, packages, and entry points.
|
||||
|
||||
Large edit: replacing `Symbol.safeDelete` implementation shape. With content-anchored Edit, a whole-body replacement would send old body plus new body, about 2x the new method payload. `replace_symbol_body` sent only the new body plus `Symbol/safeDelete`.
|
||||
Built-in call chain:
|
||||
|
||||
**Verdict:** Use text Edit for 1-3 line changes, either tool for 10-30 line method rewrites, and Serena for whole-method replacements.
|
||||
1. `git status --short` -> clean.
|
||||
2. `rg --files -g '!*.md' -g '!docs/**' -g '!CLAUDE.md'` -> code/config/resource inventory.
|
||||
3. `Get-ChildItem -Force` -> top-level directories.
|
||||
4. `Get-Content src/main/resources/META-INF/plugin.xml` -> plugin entry points.
|
||||
5. `Get-Content build.gradle.kts -TotalCount 80` -> Gradle/IntelliJ platform setup.
|
||||
|
||||
### 3.4 Refactors
|
||||
Serena call chain: not applicable for repository inventory and non-code config.
|
||||
|
||||
Private rename: `Logger/logToToolWindow` to `appendToToolWindow` changed 7 occurrences in 1 semantic call plus optional `rg` verification. Manual path was `rg`, patch each occurrence, then `rg` verify: 3 calls and more payload.
|
||||
Observed structure:
|
||||
|
||||
Multi-file rename: `Logger` to `SerenaLogger` changed 10 files and renamed the Java file. Manual path would be `rg`, read 10 files, rename/move file, edit imports/type names/constructors, verify with `rg`, and likely compile: roughly 14-20 calls.
|
||||
- `service`: backend service and request handling.
|
||||
- `service/endpoint`: endpoint handlers for symbol search, references, formatting, rename, move, safe delete, inline, inspections, and completions.
|
||||
- `symbol`: symbol model, symbol lookup, hierarchy, path matching, and move processors.
|
||||
- `util`: IDE/project/editor/UI helpers.
|
||||
- `ui`: tool window content.
|
||||
- `plugin.xml`: registers `PluginStartupActivity`, settings service/configurable, and a dummy test action.
|
||||
|
||||
Move: moving `ProjectUtil` to `service.endpoint` moved the file, changed its package, and removed the import from `RefreshFileHandler` in 1 call. Manual path would coordinate filesystem move, package line, import removal/additions, and verification.
|
||||
Payloads: built-ins used 5 small shell/read calls and returned several KB of code/config context; Serena had no relevant semantic operation.
|
||||
|
||||
Safe delete: `DebugUtil` had no references; safe delete returned `affected_references: []` and deleted it. Manual path is search, delete, verify. Inline: `ProjectUtil/getAbsolutePath` inlined into both call sites in 1 call.
|
||||
**Verdict:** Repository layout and non-code entry-point discovery are built-in work; Serena starts adding value once the target is code symbols.
|
||||
|
||||
**Verdict:** Every multi-file or semantic refactor tested was a clear Serena win in call count, payload, and correctness surface.
|
||||
### 3.2 Large file structural overview
|
||||
|
||||
### 3.5 Session Effects
|
||||
Attempted: compare structural overview on `src/main/java/de/oraios/serena/symbol/Symbol.java`.
|
||||
|
||||
I chained three edits in `Logger.java`: rename helper, replace `Logger/warning[0]`, insert after `Logger/error[0]`. The name paths survived earlier edits; no line recalculation was needed. A built-in line/slice workflow would need refreshed context after insertions because line numbers and nearby anchors shift.
|
||||
Serena call chain:
|
||||
|
||||
**Verdict:** Serena's stable identifiers compound across a session; the more edits you make in one file, the wider the gap gets.
|
||||
1. `get_symbols_overview(relative_path=Symbol.java, depth=1)`
|
||||
2. Next step: `find_symbol(Symbol/findReferences, include_body=true)`
|
||||
|
||||
## 4. Token Efficiency
|
||||
Serena output: one top-level `Symbol` class; fields; nested `ChildrenCollector` and `DocumentationResolver`; methods including overload-indexed names such as `getLocationString[0]`, `getLocationString[1]`, `move[0]`, `move[1]`, `getDocumentation[0]`, and `getDocumentation[1]`.
|
||||
|
||||
The crossover is size-dependent. Small edit: text wins because it sends only a tiny anchor. Medium rewrite: near parity. Whole-body rewrite: Serena wins because it sends new body only, while content-anchored Edit sends old body plus new body.
|
||||
Built-in call chain:
|
||||
|
||||
Forced reads matter: `find_symbol(...include_body=true)` avoided reading 573 lines of `UIControlUtil.java` and 1,295 lines of `Symbol.java`. Stable outputs also have longer shelf life: `Logger/warning[0]` remains useful after unrelated edits; line 52 does not.
|
||||
1. `rg` for method-like declaration lines in `Symbol.java`.
|
||||
2. `rg` for class/interface/enum lines in `Symbol.java`.
|
||||
3. Next step: line-slice read around the selected method.
|
||||
|
||||
**Verdict:** Token economics favor built-ins for tiny local text edits and Serena for symbol retrieval, chained work, and whole-symbol replacement.
|
||||
Built-in output: about 70 line-oriented method/class matches, including nested symbols, with no durable symbol identity beyond manual signature inspection.
|
||||
|
||||
## 5. Reliability
|
||||
Payloads: Serena used 1 compact overview call plus a targeted follow-up. Built-ins used 2 regex searches plus a later line-range read.
|
||||
|
||||
Semantic matching distinguished overloads: `Logger/warning[0]` was `warning(String,Object...)`; `Logger/warning[1]` was `warning(String,Throwable)`. `rg "warning\\("` returned both and left disambiguation to the caller.
|
||||
**Verdict:** Serena's overview is more actionable because its output feeds directly into stable symbol-addressed calls; regex outlines are useful but line-based.
|
||||
|
||||
Atomicity matters: semantic rename/move run as IDE refactorings, so a failure is not a half-finished sequence of 10 manual edits. Both toolsets still report mechanical success only; semantic intent still needs diff/build review.
|
||||
### 3.3 Targeted method body retrieval
|
||||
|
||||
Transitive queries are where text cannot compete directly: hierarchy and external declaration lookup depend on IDE indexes and dependency sources, not just strings.
|
||||
Attempted: retrieve `Symbol/findReferences` without reading surrounding file content.
|
||||
|
||||
**Verdict:** Correctness weight favors Serena for code-entity scope and refactoring atomicity, while text remains correct for text questions.
|
||||
Serena call chain: `find_symbol(relative_path=Symbol.java, name_path_pattern=Symbol/findReferences, include_body=true)`.
|
||||
|
||||
## 6. Workflow Effects
|
||||
Serena output: only the `public ArrayList<SymbolReference> findReferences()` body.
|
||||
|
||||
Over a multi-step session, Serena's intermediate artifacts stay usable: symbol trees, name paths, reference lists, and hierarchy nodes survive edits outside those symbols. Text outputs are often ephemeral: line numbers and byte offsets decay immediately after insertions/deletions.
|
||||
Built-in call chain:
|
||||
|
||||
The practical effect is not just fewer calls; it is fewer re-reads. In the chained `Logger.java` session, Serena needed three edit calls. A text workflow would normally require initial reads/searches, edits, and refreshed context before later insertions.
|
||||
1. `rg "findReferences\\(" src/main/java -n -C 2`
|
||||
2. `Get-Content Symbol.java` line slice around the declaration.
|
||||
|
||||
**Verdict:** Serena's session-level multiplier comes from not having to rediscover where code moved after each edit.
|
||||
Built-in output: search context across multiple files, then the selected method slice.
|
||||
|
||||
## 7. No Built-In Equivalent
|
||||
Payloads: Serena input was one path/name-path request and returned roughly 25 method lines. Built-ins required search output plus a separate slice read.
|
||||
|
||||
- IDE semantic rename/move/inline/safe-delete: high value when refactoring, moderate frequency. Built-ins can approximate with many edits but cannot provide IDE refactoring semantics or atomicity.
|
||||
- Type hierarchy including external `Object`: medium value, occasional. Text can search `extends`, but cannot transitively resolve hierarchy with dependency/stub awareness in one call.
|
||||
- External dependency declaration resolution: medium value, occasional. Built-ins require build-file discovery and local source/cache spelunking.
|
||||
- Overload/name-path targeting: high value in typed code, frequent enough to matter. Text can match names but cannot address `warning[0]` as a distinct method without manual signature reasoning.
|
||||
**Verdict:** If the symbol is known, Serena retrieves the body in one precise call; built-ins need search plus line-targeted reading.
|
||||
|
||||
**Verdict:** Serena's unique capabilities are concentrated in IDE-indexed code intelligence and refactoring operations, not general file manipulation.
|
||||
### 3.4 References: code usages vs mentions anywhere
|
||||
|
||||
## 8. Built-Ins As Default
|
||||
Attempted: find all references to `Symbol/findReferences`.
|
||||
|
||||
Use built-ins for non-code files, config, docs, and free text. `build.gradle.kts` was best read directly. `rg` was clearly right for `/findSymbol`, `127.0.0.1`, and `FORM_INIT_DELAY_MILLIS`.
|
||||
Serena call chain: `find_referencing_symbols(relative_path=Symbol.java, name_path=Symbol/findReferences)`.
|
||||
|
||||
Use built-ins for tiny edits where a short unique anchor is obvious. Also always keep a post-refactor text sweep for comments, markdown, notebooks, generated files, and product strings; that is complementary verification, not a Serena weakness.
|
||||
Serena output:
|
||||
|
||||
Estimated share: built-ins remain best for maybe 40-60% of daily interactions because much coding work is still file/text/config/search. Serena dominates the code-symbol subset.
|
||||
- `Symbol/formatReferenceLocations/references`
|
||||
- `Symbol/inline/refCountBefore`
|
||||
- `Symbol/verifyInlineResult/refCountAfter`
|
||||
- `SymbolDTO/Builder/buildDTO/references`
|
||||
- `RunInspectionsOnSymbolsHandler/handleRequest/refs`
|
||||
- `FindReferencesHandler/buildResponse/references`
|
||||
|
||||
**Verdict:** Start with built-ins for text and config; switch to Serena as soon as the noun in your task is a symbol.
|
||||
Built-in call chain: `rg "findReferences" src/main/java -n -C 1`.
|
||||
|
||||
## 9. Practical Rule
|
||||
Built-in output: same call sites plus the declaration, `/findReferences` route registration, and explanatory comments.
|
||||
|
||||
Reach for Serena when the task says class, method, overload, implementation, reference, hierarchy, rename, move, inline, delete, or "insert after this method." Reach for `rg`/Read/Edit when the task says string, config, docs, log message, URL, generated text, or "change these two lines."
|
||||
Payloads: Serena returned a compact grouped usage list of about 500-700 characters. `rg` returned broader line context of about 1-2 KB and mixed code usages with text mentions.
|
||||
|
||||
For edits: text Edit for 1-3 line tweaks; Serena `replace_symbol_body` for full methods/classes; semantic refactor tools for any rename/move/delete/inline that crosses call sites or imports.
|
||||
**Verdict:** Serena has higher precision for "who uses this symbol in code"; built-ins have higher recall for "where is this text mentioned anywhere."
|
||||
|
||||
I restored all tracked edits after the experiments. The tracked working tree was clean after evaluation; the only baseline untracked entries were `.claude/` and `serena-evaluation-prompt.md`. I did not run the Gradle test suite because the task was an evaluation, not a product change.
|
||||
### 3.5 Type hierarchy
|
||||
|
||||
**Verdict:** With both toolsets installed, use built-ins for text and Serena for code entities; that rule captures almost all of the measured value without overthinking each call.
|
||||
Attempted: list supertypes and subtypes transitively for `PostRequestHandler`.
|
||||
|
||||
Serena call chain: `type_hierarchy(relative_path=PostRequestHandler.java, name_path=PostRequestHandler, hierarchy_type=both, depth=0)`.
|
||||
|
||||
Serena output: external `Object` supertype; 18 direct endpoint subclasses; transitive nested branch `TypeHierarchyHandler` containing `GetSupertypesHandler` and `GetSubtypesHandler`.
|
||||
|
||||
Built-in call chain:
|
||||
|
||||
1. `rg "extends PostRequestHandler|extends TypeHierarchyHandler|class PostRequestHandler" src/main/java -n`
|
||||
2. Manually follow any discovered intermediate types.
|
||||
3. Repeat searches if deeper hierarchy exists.
|
||||
|
||||
Payloads: Serena used 1 hierarchy query and returned structured JSON. Built-ins used pattern search and manual transitive grouping.
|
||||
|
||||
**Verdict:** Serena turns hierarchy discovery into a semantic graph query; text search requires iterative pattern expansion and manual reasoning.
|
||||
|
||||
### 3.6 External dependency symbol lookup
|
||||
|
||||
Attempted: retrieve the IntelliJ dependency symbol behind `ReferencesSearch.search(anchorElement)`.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `find_declaration(relative_path=Symbol.java, regex="ReferencesSearch\\.(search)\\(anchorElement\\)")`
|
||||
2. `find_symbol(relative_path=<ext:ReferencesSearch.class|466808a0>, name_path_pattern=ReferencesSearch/search[0], include_body=true, search_deps=true)`
|
||||
|
||||
Serena output: declaration `ReferencesSearch/search[0]` in an external class path and the overload body:
|
||||
|
||||
```java
|
||||
public static @NotNull Query<PsiReference> search(@NotNull PsiElement element) {
|
||||
return search(element, GlobalSearchScope.allScope(PsiUtilCore.getProjectInReadAction(element)), false);
|
||||
}
|
||||
```
|
||||
|
||||
Built-in call chain:
|
||||
|
||||
1. `rg "ReferencesSearch" src/main/java -n` -> import and local call.
|
||||
2. `rg` over repo, `.intellijPlatform`, and Gradle caches for `class ReferencesSearch` or matching method signatures -> no useful source hit.
|
||||
3. `Get-ChildItem` over Gradle/IntelliJ caches -> jar candidates but no direct definition.
|
||||
|
||||
Infrastructure difference: Serena depends on the JetBrains IDE index; built-ins need source jars, decompiler tooling, or classpath-specific jar inspection.
|
||||
|
||||
**Verdict:** External dependency lookup is a genuine Serena capability when IDE indexes are available; ordinary built-ins do not provide it.
|
||||
|
||||
### 3.7 Single-file edits across edit sizes
|
||||
|
||||
Small tweak attempted: change `"No symbol found for "` to `"No matching symbol found for "` in `SymbolFinder/findSymbolByNamePath`.
|
||||
|
||||
Built-in call chain:
|
||||
|
||||
1. Use existing search/body context.
|
||||
2. `apply_patch` one-line replacement.
|
||||
3. `git diff`, `git status --short`.
|
||||
4. Revert and clean check.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `replace_symbol_body(SymbolFinder/findSymbolByNamePath, full method body with changed string)`.
|
||||
2. `git diff`, `git status --short`.
|
||||
3. Revert and clean check.
|
||||
|
||||
Observed diff: both produced the same 1-line change. Payload difference: built-in edit input was a tiny hunk; Serena input was the full 11-line method body.
|
||||
|
||||
Medium rewrite attempted: rewrite `SymbolFinder/findFilesByName` from list accumulation to stream collection.
|
||||
|
||||
Built-in call chain: `apply_patch` over the method body, then diff/status/revert/status.
|
||||
|
||||
Serena call chain: `replace_symbol_body(SymbolFinder/findFilesByName, new method body)`, then diff/status/revert/status.
|
||||
|
||||
Observed diff: both produced `10 +++-------`, 3 insertions and 7 deletions. Payload difference was small: a patch hunk versus an 8-line replacement method.
|
||||
|
||||
Large/whole-body rewrite attempted: rewrite `InspectionRunner/collectResults` while preserving signature and behavior shape.
|
||||
|
||||
Built-in call chain: `apply_patch` against the method body, then diff/status/revert/status.
|
||||
|
||||
Serena call chain: `replace_symbol_body(InspectionRunner/collectResults, full symbol text including Javadoc and replacement body)`, then diff/status/revert/status.
|
||||
|
||||
Observed diff: both produced `38 +++++++++++-----------`, 19 insertions and 19 deletions. Payload difference: built-in required a large contextual hunk; Serena required the full symbol text, about 70 lines including Javadoc/signature/body.
|
||||
|
||||
**Verdict:** Serena is not automatically more token-efficient for edits; its edit advantage is stable symbol targeting, while built-ins are smaller for tiny local hunks.
|
||||
|
||||
### 3.8 Insert method at structural location
|
||||
|
||||
Attempted: insert `private Logger getLogger()` immediately after `SymbolFinder/getProject`.
|
||||
|
||||
Built-in call chain:
|
||||
|
||||
1. Search/read surrounding area.
|
||||
2. `apply_patch` with nearby context.
|
||||
3. Diff/status/revert/status.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `insert_after_symbol(relative_path=SymbolFinder.java, name_path=SymbolFinder/getProject, body=...)`
|
||||
2. Diff/status/revert/status.
|
||||
|
||||
Observed diff: both produced the same 4-line insertion.
|
||||
|
||||
Payloads: built-ins sent inserted body plus surrounding text context; Serena sent inserted body plus symbol target.
|
||||
|
||||
**Verdict:** Structural insertion is a modest Serena improvement: same resulting diff, less dependence on line/context stability.
|
||||
|
||||
### 3.9 Rename: private helper in one file
|
||||
|
||||
Attempted: rename `SymbolFinder/qualNameMatchesName` to `qualifiedNameMatchesName`.
|
||||
|
||||
Built-in call chain:
|
||||
|
||||
1. `rg "qualNameMatchesName" SymbolFinder.java -n -C 2`
|
||||
2. `apply_patch` declaration and one call site.
|
||||
3. Diff/status/revert/status.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `find_referencing_symbols(SymbolFinder/qualNameMatchesName)` -> one referencing method.
|
||||
2. `rename(name_path=SymbolFinder/qualNameMatchesName, new_name=qualifiedNameMatchesName, rename_in_comments=true, rename_in_text_occurrences=true)`.
|
||||
3. Diff/status/revert/status.
|
||||
|
||||
Observed result: Serena returned `"Success"`. Both produced the same 2-line diff.
|
||||
|
||||
**Verdict:** Semantic rename has little advantage for a one-file helper with one call site, but it scales better than text edits as references spread or names become ambiguous.
|
||||
|
||||
### 3.10 Rename: cross-file class including imports
|
||||
|
||||
Attempted: rename `MoveOperationFailedException` to `MoveFailedException`.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `rename(relative_path=MoveOperationFailedException.java, name_path=MoveOperationFailedException, new_name=MoveFailedException, rename_in_comments=true, rename_in_text_occurrences=true)`.
|
||||
2. `git diff --stat`, `git diff`, `git diff --cached --stat`, `git status --short`.
|
||||
3. Restore staged rename and modified files; clean check.
|
||||
|
||||
Serena result: `"Success"`.
|
||||
|
||||
Observed edits:
|
||||
|
||||
- Added/staged `MoveFailedException.java`.
|
||||
- Deleted/staged `MoveOperationFailedException.java`.
|
||||
- Updated import, Javadoc comment, and thrown class in `MoveProcessor.java`.
|
||||
|
||||
Built-in equivalent chain:
|
||||
|
||||
1. `rg "MoveOperationFailedException" src/main/java -n`.
|
||||
2. Rename/create/delete file.
|
||||
3. Patch class name and constructor.
|
||||
4. Patch imports and usages.
|
||||
5. Patch comments/text occurrences if desired.
|
||||
6. Verify no unintended old references remain.
|
||||
7. Compile/test for a permanent change.
|
||||
|
||||
**Verdict:** Cross-file class renames are high-value Serena cases because file/class coupling, imports, usages, and optional comments are handled as one semantic refactor.
|
||||
|
||||
### 3.11 Move symbol to another file context
|
||||
|
||||
Attempted: move nested record `InspectionRunner/PendingFix` to a top-level file in the same package.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `move(relative_path=InspectionRunner.java, name_path=InspectionRunner/PendingFix, target_relative_path=src/main/java/de/oraios/serena/service/endpoint)`.
|
||||
2. Diff/status/cached-diff inspection.
|
||||
3. Restore staged new file and modified files; clean check.
|
||||
|
||||
Serena result:
|
||||
|
||||
```json
|
||||
{
|
||||
"source_relative_path": "src/main/java/de/oraios/serena/service/endpoint/InspectionRunner.java",
|
||||
"target_relative_path": "src/main/java/de/oraios/serena/service/endpoint",
|
||||
"moved_symbol": {"name_path": "PendingFix", "type": "CLASS"}
|
||||
}
|
||||
```
|
||||
|
||||
Observed edits:
|
||||
|
||||
- Created/staged `PendingFix.java` with package and imports.
|
||||
- Removed nested record from `InspectionRunner.java`.
|
||||
- Updated 3 external references from `InspectionRunner.PendingFix` to `PendingFix`.
|
||||
- Removed unused imports in 2 files.
|
||||
|
||||
Built-in equivalent chain: read nested record/import needs, add file, delete nested symbol, search usages, patch qualified usages, patch internal references if needed, remove imports, verify, compile/test.
|
||||
|
||||
**Verdict:** Symbol move is a substantial Serena addition because the difficult part is reference/import repair, not copying text.
|
||||
|
||||
### 3.12 Move file/package location
|
||||
|
||||
Attempted: move `PluginUtil.java` from `util` to `service`.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `move(relative_path=src/main/java/de/oraios/serena/util/PluginUtil.java, target_relative_path=src/main/java/de/oraios/serena/service)`.
|
||||
2. Diff/status/cached-diff inspection.
|
||||
3. Restore staged rename and modified imports; clean check.
|
||||
|
||||
Serena result:
|
||||
|
||||
```json
|
||||
{
|
||||
"source_relative_path": "src/main/java/de/oraios/serena/util/PluginUtil.java",
|
||||
"target_relative_path": "src/main/java/de/oraios/serena/service"
|
||||
}
|
||||
```
|
||||
|
||||
Observed edits:
|
||||
|
||||
- Staged rename `util/PluginUtil.java -> service/PluginUtil.java`.
|
||||
- Changed package declaration to `de.oraios.serena.service`.
|
||||
- Updated import in `PluginStartupActivity`.
|
||||
- Removed now-unneeded same-package import in `SerenaBackendService`.
|
||||
|
||||
Built-in equivalent chain: move file, patch package declaration, search `PluginUtil`, patch imports/usages, remove redundant imports, verify no old package import remains, compile/test.
|
||||
|
||||
**Verdict:** File moves are high-value when packages/imports must change; built-ins can do them manually but require dependency repair steps.
|
||||
|
||||
### 3.13 Safe delete and propagated delete
|
||||
|
||||
Attempted: delete unused `DebugUtil`.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `safe_delete(relative_path=DebugUtil.java, name_path=DebugUtil, delete_even_if_used=false, propagate=false)`.
|
||||
2. `git diff --cached --stat`, `git status --short`.
|
||||
3. Restore staged deletion; clean check.
|
||||
|
||||
Serena result:
|
||||
|
||||
```json
|
||||
{
|
||||
"deleted_symbol": "DebugUtil",
|
||||
"relative_path": "src/main/java/de/oraios/serena/util/DebugUtil.java",
|
||||
"affected_references": [],
|
||||
"message": "Symbol deleted successfully with no affected references."
|
||||
}
|
||||
```
|
||||
|
||||
Built-in equivalent chain: `rg "DebugUtil"`, inspect declaration vs usages, delete file/symbol, verify no remaining usages.
|
||||
|
||||
Propagated deletion: no suitable correct-use candidate was found where JetBrains safe-delete propagation would meaningfully delete a used symbol and propagate deletion through call sites. For ordinary used-method deletion, correct safe delete reports usages or, if forced, deletes the symbol and reports affected references; that is not the same as arbitrary call-site deletion.
|
||||
|
||||
**Verdict:** Safe delete adds an integrated usage check plus deletion operation; propagated deletion was not evaluated because this codebase did not provide a suitable candidate.
|
||||
|
||||
### 3.14 Inline
|
||||
|
||||
Attempted: inline `ProjectUtil/getAbsolutePath`, a single-expression helper.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `find_symbol(ProjectUtil/getAbsolutePath, include_body=true)`.
|
||||
2. `find_referencing_symbols(ProjectUtil/getAbsolutePath)` -> 2 references.
|
||||
3. `inline_symbol(relative_path=ProjectUtil.java, name_path=ProjectUtil/getAbsolutePath, keep_definition=false)`.
|
||||
4. Diff/status/revert/status.
|
||||
|
||||
Serena result: `{"status": "SUCCESS"}`.
|
||||
|
||||
Observed edits:
|
||||
|
||||
- Removed `getAbsolutePath`.
|
||||
- Replaced internal call in `ProjectUtil/getVirtualFile`.
|
||||
- Replaced external call in `RefreshFileHandler`.
|
||||
|
||||
Built-in equivalent chain: read helper body, `rg getAbsolutePath`, patch each call site with substituted expression, delete helper, verify no stale references, compile/test.
|
||||
|
||||
**Verdict:** Inline is a strong Serena refactor for legally substitutable helpers; built-ins can reproduce it but must manually adapt each substitution.
|
||||
|
||||
### 3.15 Scope precision and overload targeting
|
||||
|
||||
Attempted: distinguish overloaded `Symbol/getLocationString` methods.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `find_symbol(Symbol/getLocationString[0], include_body=true)`.
|
||||
2. `find_symbol(Symbol/getLocationString[1], include_body=true)`.
|
||||
|
||||
Serena output:
|
||||
|
||||
- `[0]`: `public String getLocationString() { return getLocationString(element); }`
|
||||
- `[1]`: `private static String getLocationString(PsiElement element) { ... }`
|
||||
|
||||
Built-in call chain: `rg "getLocationString" Symbol.java -n -C 1`, then manual signature inspection.
|
||||
|
||||
**Verdict:** Serena's overload-indexed name paths provide precise symbol addressing that text search does not.
|
||||
|
||||
### 3.16 Atomicity and success signals
|
||||
|
||||
Observed Serena success signals:
|
||||
|
||||
- `replace_symbol_body` -> `OK`
|
||||
- `insert_after_symbol` -> `OK`
|
||||
- `rename` -> `"Success"`
|
||||
- `move` -> JSON with source/target/moved symbol or paths
|
||||
- `safe_delete` -> JSON with deleted symbol, affected references, and message
|
||||
- `inline_symbol` -> `{"status": "SUCCESS"}`
|
||||
|
||||
Built-in success signals:
|
||||
|
||||
- `apply_patch` -> "Success. Updated the following files"
|
||||
- Shell/Git commands -> exit codes and textual output
|
||||
- Cross-file consistency -> user-driven diff/search/build verification
|
||||
|
||||
Atomicity comparison: Serena refactors executed as single IDE operations. Built-in equivalents are chains of searches, file operations, and patches; a competent user can make them correct, but intermediate states can be partial.
|
||||
|
||||
Tradeoff: class/file move/delete refactors staged added/deleted files in the Git index, so clean verification must check both staged and unstaged state.
|
||||
|
||||
**Verdict:** Serena gives clearer semantic success signals and more atomic cross-file edits; built-ins require manual consistency management.
|
||||
|
||||
### 3.17 Workflow effects across multiple edits
|
||||
|
||||
Attempted: chain 3 edits in `SymbolFinder.java` without refreshing between them.
|
||||
|
||||
Serena call chain:
|
||||
|
||||
1. `replace_symbol_body(SymbolFinder/findFilesByName, ...)`.
|
||||
2. `insert_after_symbol(SymbolFinder/getProject, ...)`.
|
||||
3. `rename(SymbolFinder/qualNameMatchesName, ...)`.
|
||||
4. Diff/status/revert/status.
|
||||
|
||||
Observed result: all edits applied after earlier edits shifted file positions. Final diff touched 18 lines with net 9 insertions and 9 deletions.
|
||||
|
||||
Built-in equivalent chain: initial read/search, apply first patch, rely on robust patch context or refresh line numbers, apply second patch, refresh or search again, apply rename patch, verify.
|
||||
|
||||
**Verdict:** Serena's name paths remain useful after line shifts, so its advantage compounds in multi-edit sessions.
|
||||
|
||||
### 3.18 Non-code reads and free-text search
|
||||
|
||||
Attempted: read config and search free text.
|
||||
|
||||
Built-in call chain:
|
||||
|
||||
1. `Get-Content plugin.xml`.
|
||||
2. `Get-Content build.gradle.kts -TotalCount 80`.
|
||||
3. `rg "http://|https://|localhost|127\\.0\\.0\\.1|8080|24282|SERENA|serena" src/main/java src/main/resources -n`.
|
||||
|
||||
Serena call chain: not applicable.
|
||||
|
||||
Observed result: built-ins exposed plugin registration, Gradle platform setup, URLs, IDs, settings defaults, and resource/text mentions.
|
||||
|
||||
**Verdict:** Built-ins are natural and sufficient for config reading and free-text search; this is outside Serena's semantic-code scope.
|
||||
|
||||
## 4. Token-efficiency analysis
|
||||
|
||||
Payload differences across edit sizes:
|
||||
|
||||
| Task | Built-in payload | Serena payload | More efficient |
|
||||
|---|---|---|---|
|
||||
| 1-line string tweak | Tiny patch hunk, 1 changed line | Full 11-line method body | Built-ins |
|
||||
| Medium method rewrite | 10-line diff hunk | Full 8-line replacement method | Roughly equal |
|
||||
| Large body rewrite | Large contextual patch, 38-line diff | Full symbol text, about 70 lines | Roughly equal; Serena has better target |
|
||||
| Insert method | Inserted body plus surrounding context | Inserted body plus name path | Serena slightly |
|
||||
| Private one-file rename | `rg` plus 2-line patch | optional refs query plus rename command | Roughly equal |
|
||||
| Cross-file rename/move/inline | Multiple searches, file operations, patches, verifications | One semantic refactor plus diff/status | Serena |
|
||||
|
||||
Forced reads:
|
||||
|
||||
- Built-ins commonly need a search before a precise read, then a line slice or full context before editing.
|
||||
- Serena can skip full-file reads when the symbol name path is known.
|
||||
- When the symbol is not known, Serena's `get_symbols_overview` and shallow `find_symbol` calls provide compact discovery rather than full-file reads.
|
||||
|
||||
Output payload:
|
||||
|
||||
- Serena exploration output is structured and scoped to symbols.
|
||||
- Built-in grep output can be larger because it includes declarations, comments, strings, and unrelated text mentions.
|
||||
- Built-ins can be extremely low-output for known-location small edits.
|
||||
|
||||
Stable vs ephemeral addressing:
|
||||
|
||||
- Built-ins address code by file paths, line numbers, text context, or byte positions. Line numbers and slices go stale after edits.
|
||||
- Serena addresses symbols by `relative_path + name_path`, with overload indexes where needed. The chained-edit experiment showed those targets survived line shifts.
|
||||
- Path-changing operations still require updated paths afterward, so symbol stability is strongest before moves/renames that alter file locations.
|
||||
|
||||
**Verdict:** Serena saves tokens by avoiding broad reads/search triage and by collapsing multi-file edit chains; built-ins remain more token-efficient for tiny known-location hunks.
|
||||
|
||||
## 5. Reliability & correctness under correct use
|
||||
|
||||
Precision of matching:
|
||||
|
||||
- Serena reference search returned code usages and excluded route strings/comments.
|
||||
- Built-in text search returned all mentions and therefore mixed true code usage with broader text hits.
|
||||
- Serena overload indexes targeted `Symbol/getLocationString[0]` and `[1]`; text search required manual signature inspection.
|
||||
|
||||
Scope disambiguation:
|
||||
|
||||
- Serena name paths encode class nesting and overload identity.
|
||||
- Built-ins can disambiguate with careful patterns and reads, but the disambiguation is manual.
|
||||
- In simple cases, such as a private helper with one call site, this semantic precision had little practical value. In overloaded, nested, or cross-file cases, it materially reduced ambiguity.
|
||||
|
||||
Atomicity:
|
||||
|
||||
- Serena refactors apply through JetBrains refactoring operations and produce one operation-level success signal.
|
||||
- Built-in equivalents are decomposed into separate searches, file moves, patches, and verification steps.
|
||||
- Competent users can verify either path, but Serena reduces the number of intermediate inconsistent states.
|
||||
|
||||
Semantic queries vs text search:
|
||||
|
||||
- `type_hierarchy` and dependency declaration lookup produced semantic information not available from ordinary grep.
|
||||
- `rg` remained better for broad mention searches and config/resource scans.
|
||||
|
||||
External dependency limitations:
|
||||
|
||||
- Serena's dependency lookup depends on IDE/language indexes and available dependency metadata.
|
||||
- Built-ins can match or exceed it only if paired with source jars, decompilers, or separate language-server tooling.
|
||||
|
||||
Operational tradeoff:
|
||||
|
||||
- Several Serena refactors staged created/deleted files. This is not a correctness problem, but it adds index-state verification to the cleanup workflow.
|
||||
|
||||
**Verdict:** Serena improves correctness where symbol identity and cross-file semantics matter; built-ins remain reliable and simpler for text/file tasks.
|
||||
|
||||
## 6. Workflow effects across a session
|
||||
|
||||
Where Serena advantages compound:
|
||||
|
||||
- Symbol overview results become inputs to body reads, reference searches, renames, moves, safe deletes, and inline operations.
|
||||
- Name paths remain useful after line shifts, reducing re-read/re-target work during multi-edit sessions.
|
||||
- Refactor-heavy sessions benefit because one semantic command replaces search/edit/verify loops across files.
|
||||
|
||||
Where Serena advantages diminish:
|
||||
|
||||
- After a full file has already been read, overview adds less incremental value.
|
||||
- If the task is a small textual substitution in a known place, symbolic body replacement can cost more payload than a tiny patch.
|
||||
- For non-code and free-text work, Serena has no role.
|
||||
|
||||
Intermediate result durability:
|
||||
|
||||
- Serena intermediate results such as `SymbolFinder/findFilesByName` or `Symbol/getLocationString[1]` remain meaningful as long as the file/symbol still exists.
|
||||
- Built-in line numbers became stale after insertions; robust textual patch context remained usable when surrounding lines did not change.
|
||||
|
||||
Verification cost:
|
||||
|
||||
- Both workflows still need `git diff`, build/test, and domain-specific verification for permanent changes.
|
||||
- Serena success signals reduce the need to inspect every call site manually, but not the need to inspect the intended diff.
|
||||
|
||||
**Verdict:** Serena's advantages compound inside symbol-centric sessions and diminish once the work becomes plain text, config, shell, or already-loaded single-file editing.
|
||||
|
||||
## 7. Unique capabilities
|
||||
|
||||
Unique here means no practical built-in equivalent without adding separate language-server, IDE, source-index, or decompiler infrastructure.
|
||||
|
||||
- Transitive type hierarchy. Frequency: medium. Impact: high when changing base classes, interfaces, or handler contracts. Built-ins can grep explicit `extends` clauses but do not produce a semantic transitive graph or external supertypes.
|
||||
|
||||
- External dependency declaration/body lookup from a call site. Frequency: low-medium. Impact: high for API behavior questions. Serena resolved `ReferencesSearch.search` into an external class method body; ordinary built-ins did not.
|
||||
|
||||
- IDE semantic move of a nested symbol to a top-level file with usage/import repair. Frequency: low-medium. Impact: high during refactors. Built-ins can reproduce this only as a manual multi-step refactor.
|
||||
|
||||
- Safe delete with semantic usage check. Frequency: medium. Impact: medium-high. Built-ins can search and delete, but they do not provide an integrated safe-delete operation over code usages.
|
||||
|
||||
- Inline refactor with call-site substitution. Frequency: low-medium. Impact: medium-high. Built-ins can patch substitutions manually, but expression adaptation and helper deletion are not primitive text operations.
|
||||
|
||||
Not unique: reading files, listing files, grep, small patches, simple one-file renames, config/resource understanding, and shell/Git operations.
|
||||
|
||||
**Verdict:** Serena's unique practical capabilities are hierarchy/dependency queries and IDE refactors; its non-unique areas are ordinary text and file operations.
|
||||
|
||||
## 8. Tasks outside Serena's scope
|
||||
|
||||
Built-in-only or built-in-natural tasks observed:
|
||||
|
||||
- Repository inventory and top-level layout.
|
||||
- Reading `plugin.xml`, `build.gradle.kts`, scripts, resources, and generated artifacts.
|
||||
- Free-text search for URLs, settings strings, endpoint strings, magic constants, resource IDs, and comments.
|
||||
- Running shell commands, builds, tests, Git operations, and filesystem cleanup.
|
||||
- Small text edits where a one-line patch is the whole task.
|
||||
|
||||
Estimated share of daily coding work:
|
||||
|
||||
- About 30-50% of a normal coding session is built-in-natural: shell, tests, grep, config reads, small text patches, and Git checks.
|
||||
- Serena applies to much of the remaining 50-70% when the work is code navigation, symbol understanding, references, hierarchy, or refactoring.
|
||||
- Serena coverage rises in refactor-heavy sessions and falls in config/debug/test-log sessions.
|
||||
|
||||
**Verdict:** Serena augments the semantic-code portion of development; built-ins still carry a large and necessary share of routine work.
|
||||
|
||||
## 9. Practical usage rule
|
||||
|
||||
Use Serena when the task can be stated as a code-symbol operation:
|
||||
|
||||
- "What methods/classes are in this file?"
|
||||
- "Give me this method body."
|
||||
- "Who uses this symbol?"
|
||||
- "What implements/subclasses this?"
|
||||
- "Rename, move, delete, or inline this symbol and update usages."
|
||||
- "Find the declaration of this call, including dependency symbols."
|
||||
- "Apply several edits where line locations may shift."
|
||||
|
||||
Use built-ins when the task is text, file, shell, config, or tiny local editing:
|
||||
|
||||
- "List files."
|
||||
- "Read config."
|
||||
- "Search for this string anywhere."
|
||||
- "Change this one known line."
|
||||
- "Run tests/build."
|
||||
- "Inspect Git state."
|
||||
- "Edit markdown/resources/scripts."
|
||||
|
||||
For mixed workflows:
|
||||
|
||||
1. Use built-ins to find files/config and run verification.
|
||||
2. Use Serena to identify and transform code symbols.
|
||||
3. Use built-ins to inspect final diffs and run tests.
|
||||
4. Prefer built-ins for tiny local hunks after the symbol has already been located.
|
||||
5. Prefer Serena for any refactor that crosses file boundaries or depends on type/reference semantics.
|
||||
|
||||
**Verdict:** Choose Serena for symbol identity and semantic refactoring; choose built-ins for text, shell, config, verification, and minimal local edits.
|
||||
Reference in new issue
Block a user