@@ -35,22 +35,28 @@ As the ones actually applying Serena's tools, they are in the best position to e
We crafted an unbiased evaluation prompt that leads the agent to perform ~20 routine coding tasks,
representative of everyday development work,
in order to compare Serena's tools with its own built-ins, measure the differences, and report the results.
in order to estimate the value added by Serena's tools when used alongside its own built-ins.
Here's a one-sentence summary of what the agents had to say:
**Opus 4.6 (high effort) in Claude Code on a large Python codebase:**
> "Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that
**Opus 4.6 (high) in Claude Code on a large Python codebase:**
> "Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit – cross-file renames, moves, and reference lookups that
would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up."
**GPT 5.4 (high) in Codex CLI on a Java codebase:**
> "As a coding AI agent, I would ask my owner to add Serena because it gives me the missing IDE-level understanding of symbols, references, and
refactorings, turning fragile text surgery into calmer, faster, more confident code changes where semantics matter."
Give your agent the tools it has been asking for and add Serena MCP to your client!
**GPT 5.4 (medium) in Copilot CLI on a large, multi-language monorepo:**
> "As a coding agent, I’d absolutely ask my owner to add Serena because it makes me noticeably sharper and calmer on
real code – especially symbol-aware navigation, cross-file refactors, and monorepo dependency jumps – while I still lean
on built-ins for tiny text edits and non-code work."
See our [documentation](https://oraios.github.io/serena/04-evaluation/000_evaluation-intro.html) for the full methodology and much more detailed evaluation results
beyond these brief summaries, or run your own evaluation on a project of your choice.
Different agents in different settings independently converge on the same verdict.
_Give your agent the tools it has been asking for and add Serena MCP to your client!_
See our [documentation](https://oraios.github.io/serena/04-evaluation/000_evaluation-intro.html) for the full methodology and much more detailed evaluation results, or run your own evaluation on a project of your choice.
> **Claude Code (Opus 4.6):** "Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up."
> **Claude Code (Opus 4.6, medium):** "Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up."
> **Codex (GPT 5.4):** "As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and refactoring, which makes real code changes feel faster, safer, and far less blind."
> **Codex (GPT 5.4, high):** "As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and refactoring, which makes real code changes feel faster, safer, and far less blind."
These are not marketing quotes — they are the agents' own one-sentence verdicts after
performing hands-on tasks each, using both Serena's tools and their built-in
> **Codex (GPT 5.4, medium):** "As a coding agent, I’d absolutely ask my owner to add Serena because it makes me noticeably sharper and calmer on
real code – especially symbol-aware navigation, cross-file refactors, and monorepo dependency jumps – while I still lean
on built-ins for tiny text edits and non-code work."
These are not marketing quotes – they are the agents' own one-sentence verdicts after
performing hands-on tasks, using both Serena's tools and their built-in
alternatives side by side, on real codebases (large Python library in the case of Claude Code,
medium-sized Java project in the case of Codex).
Both agents independently converge on the same core finding: Serena's strongest
Different agents in different settings independently converge on the same core finding: Serena's strongest
contribution is collapsing multi-file, semantically-aware operations into single atomic
calls, while built-in tools remain the better choice for small local edits, text search,
> **One-line summary:** Serena's IDE-backed semantic tools are the single most impactful addition to my toolkit — cross-file renames, moves, and reference lookups that would cost me 8–12 careful, error-prone steps collapse into one atomic call, and I would absolutely ask any developer I work with to set them up.
> **One-line summary:** As a coding agent, I would ask my owner to add Serena because it turns fragile text-and-line-number work into precise symbol-aware navigation and
**Codebase:** [ente](https://github.com/ente-io/ente) - a large monorepo in Dart, TypeScript, Go, Rust, and other languages.
:::
# Copilot CLI (GPT-5.4, medium)
> As a coding agent, I’d absolutely ask my owner to add Serena because it makes me noticeably sharper and calmer on
real code—especially symbol-aware navigation, cross-file refactors, and monorepo dependency jumps—while I still lean
on built-ins for tiny text edits and non-code work
**Baseline.** I started from source code only, avoided repo docs/notes, and ran each reversible experiment against the repo directly. After every edit/refactor experiment, the working tree returned to its baseline state.
## 1. Headline: what Serena changes
Serena changes the workflow **when the task is about code symbols rather than raw text**. In this repo, the practical delta was:
1.**Added capability / materially better workflow.** Symbol-aware navigation and refactors in TypeScript: structural overviews, precise code-only references, hierarchy queries, symbol-targeted rename, file move with import updates, and inline. These usually collapsed a 2-6 step built-in chain into **one semantic operation after discovery**, and reduced manual scope verification.
2.**Applies but offers little or no improvement.** Small local edits inside an already-understood method. Built-ins can patch only the changed lines; Serena's body replacement resends the whole symbol, so it was often **less payload-efficient** for 1-3 line tweaks.
3.**Outside Serena's scope.** Non-code reads, free-text search, git inspection, config/package files, and other text-first tasks. Built-ins remained the natural tools there.
Two important observed limits constrained Serena's delta here: **the strongest measured gains were concentrated in the TypeScript desktop app and the Rust core crate where I ran the hands-on comparisons**, and **some refactors still carried diff-shape tradeoffs** such as formatting churn or unexpected target-file choices.
**Verdict:** In this repo, Serena was a strong TypeScript symbol layer on top of the built-ins, not a general replacement for text/file work.
## 2. Added value and differences by area
| Area | What changed vs built-ins | Frequency | Value per hit |
| --- | --- | --- | --- |
| **Cross-file symbol refactors** | `rename`, file `move`, and `inline` turned manual search/edit/update chains into one semantic op. The `wait -> delay` rename updated 4 files from 1 symbol definition; moving `http.ts` updated the importing file automatically. | Medium | High: typically **2-5 calls saved** plus less manual scope checking |
| **Code-only discovery** | Symbol overview, symbol body retrieval, reference search, and type hierarchy returned code structure directly instead of raw text matches. For `wait`, Serena returned **3 real code-use files**; `rg` returned **7 files**, including docs/comments and English-word hits. | High | Medium: usually **1-3 follow-up reads/filters avoided** |
| **Stable addressing** | Name paths stayed reusable across multiple edits (`createMainWindow`, `openStreetMapUserAgent`, `wait`); built-in line ranges had to be reacquired after edits. | Medium | Medium: less re-reading, less stale-context risk |
| **Small in-method edits** | Serena was not more efficient. Replacing `AutoLauncher/toggleAutoLaunch` required resending the full method body, while the built-in patch changed only the touched lines. | High | Low negative: built-ins used **smaller edit payloads** |
| **External dependency lookup in a monorepo** | Once indexing was available, Serena resolved Electron types from `desktop/node_modules`, `next-electron-server` declarations from the desktop package, and Rust crate symbols from Cargo registry sources. In a monorepo, that removes a manual "which package owns this dependency?" step. | Medium | Medium-High: usually **1-3 searches plus path discovery avoided** |
**Verdict:** Serena's highest-value delta was semantic refactoring plus dependency-aware code lookup in the TypeScript/Rust parts of the monorepo; its weakest area was tiny local edits.
## 3. Detailed evidence, grouped by capability
### 3.1 Codebase understanding
#### Task 1: high-level repository overview
- **Attempted:** top-level layout and likely code-heavy areas.
- **Serena chain:** `get_symbols_overview(main.ts, depth=1)` -> concise symbol map of top-level functions and nested locals under `main`; next step `find_symbol(createMainWindow, include_body=true)`.
- **Built-in chain:** `rg` on `const|function|class|export` in `main.ts` -> flat text hits; next step `view` of lines `331-439` to read `createMainWindow`.
- **Payloads observed:**
- Serena overview output: compact symbol list for the file; next-step body fetch returned only the selected symbol body.
- Built-in overview output: many matching lines without structure; next-step read required **~109 lines** of file content.
- **Delta:** Serena's overview was not just shorter; it also supplied **stable symbol names** for the follow-up call. Built-ins could answer the question, but only after a second text-localizing step.
**Verdict:** Serena materially improved the "overview -> inspect one function" flow by making the follow-up call symbol-based instead of line-based.
#### Task 3: retrieve a specific class method body without reading the surrounding file
- **Target symbol:** `AutoLauncher/toggleAutoLaunch` in `desktop/src/main/services/auto-launcher.ts`.
- **Equivalent used:** interface hierarchy in `web/apps/ensu/src/services/llm/inference.ts`, because this TS area had interface implementations rather than rich class inheritance.
- **Built-in chain:** `rg InferenceBackend|implements InferenceBackend` -> manual reconstruction from four text matches.
- **Payloads observed:** Serena returned the hierarchy directly; built-ins returned only raw declarations/usages.
- **Delta:** Serena removed the manual synthesis step. Built-ins were sufficient here because the hierarchy was shallow, but that was because the example was small.
**Verdict:** Serena added moderate value for hierarchy queries; the value grows with hierarchy depth.
#### Task 6: external dependency symbol lookup
- **Targets used after indexing was available:** `BrowserWindow` and `serveNextAt` in the desktop TypeScript app, plus `Url` and `Zeroizing` in `rust/core`.
-`find_declaration(import serveNextAt ... , include_body=true)` -> `desktop/node_modules/next-electron-server/index.d.ts`, body `declare function serveNextAt(uri: string, options?: Options): void;`
-`find_symbol(..., search_deps=true)` on those dependency files returned dependency-side docs.
- **Serena chain (Rust):**
-`find_declaration(use reqwest::{Response, Url};, include_body=true)` -> `<ext:lib.rs|...>` external symbol `Url[0]` with the struct body
-`find_declaration(use zeroize::Zeroizing;, include_body=true)` -> `<ext:lib.rs|...>` external symbol `Zeroizing[0]` with the struct body
-`find_symbol(..., relative_path=<ext...>, search_deps=true)` returned dependency-side docs for those external symbols.
- **Built-in equivalent chain:**
- Manually infer the correct monorepo-local dependency root (`desktop/node_modules`, not repo root),
- or manually inspect Cargo metadata / `Cargo.lock`,
- then open the resolved dependency files directly (for Rust, under the Cargo registry).
- **Payloads observed:** Serena returned the declaration target and a small signature/body directly; built-ins required **package-root discovery first**, which is a real extra step in a monorepo.
- **Delta:** Serena **does add capability and efficiency here** once indexing exists. The gain is larger in this monorepo than in a single-package repo because dependency ownership is split across package-local Node dependencies and shared Cargo registry sources.
**Verdict:** With indexing available, Serena added meaningful external dependency lookup, and the value was amplified by the monorepo layout.
### 3.2 Single-file edits
#### Task 7a: small tweak (1-3 lines inside a method)
- **Change:** rename local `autoLaunch` -> `launcher` inside `AutoLauncher/toggleAutoLaunch`.
- **Observed result:** Serena removed the symbol from `common.ts`, updated `ffmpeg-worker.ts`, but created a **new file**`desktop/src/main/utils/nullToUndefined.ts` instead of merging into `http.ts`.
- **Built-in equivalent:** would require manually copying the symbol into the intended target module, updating imports, then deleting the old definition.
- **Delta:** Serena still automated the cross-file update, but **did not provide the specific "move into existing module" behavior I was testing**.
**Verdict:** Serena added partial value for symbol moves here, but not the full capability of "move into a chosen existing TS file".
#### Task 12: move a file/package and update imports
- **Target:** move `desktop/src/main/utils/http.ts` to `desktop/src/main/services/http.ts`.
- **Observed result:** the file was renamed/moved and `ffmpeg-worker.ts` import updated from `../utils/http` to `./http`.
- **Built-in equivalent:** locate all imports, move the file, patch each import path, then verify.
- **Delta:** this was a real one-call semantic file move.
**Verdict:** Serena materially improved file moves that require import updates.
#### Task 12 (safe delete with no remaining usages)
- **Attempted:** searched for naturally unused TS symbols in the working areas (`main.ts`, `common.ts`, `temp.ts`, `inference.ts`) and checked several candidates (`registerForEnteLinks`, `minimumWindowSize`, `AutoLauncher/isEnabled`, `openStreetMapUserAgent`, `safeJson`, `buildSamplingConfig`).
- **Observed result:** every plausible candidate still had live references.
- **Outcome:** **no suitable candidate found in the TS areas where Serena was operational**, so I skipped this comparison instead of forcing an invalid input.
**Verdict:** No evidence either way here because the repo did not offer a clean unused-symbol candidate in the code areas Serena handled reliably.
#### Task 13: delete a symbol and propagate deletion to call sites
- **Attempted:** looked for a helper whose call sites could be semantically removed rather than inlined or manually rewritten.
- **Observed result:** the good candidates in this repo were better modeled as **inline** refactors, not delete-with-propagation.
- **Outcome:** **no suitable candidate**; skipped rather than using an unsafe input.
**Verdict:** No measured delta here because the available candidates were inline candidates, not safe propagate-delete candidates.
#### Task 13 (inline a small helper)
- **Target symbol:** `waitForRendererDevServer`.
- **Built-in chain:** `view` call site + definition -> `apply_patch` replacing `await waitForRendererDevServer()` with `await wait(1000)` and deleting the helper -> `git diff`.
- **Delta:** Serena added the unique semantic refactor, but in this run it also introduced **format churn outside the logical change**.
**Verdict:** Serena added real inline capability, with a low-frequency but real tradeoff of broader formatting churn.
### 3.4 Reliability & correctness-oriented checks
#### Task 14: scope precision
- **Demonstrated with:** `AutoLauncher/toggleAutoLaunch`, `openStreetMapUserAgent`, and `InferenceBackend`.
- **Serena:** symbol names and name paths targeted the exact code entity.
- **Built-ins:** text search for names such as `wait` or `writeToTemporaryFile` over-matched comments, docs, and multiple textual occurrences.
- **Delta:** Serena's unit of work was the symbol; built-ins' unit was the matching line.
**Verdict:** Serena was reliably more precise whenever the target was a symbol rather than a string.
#### Task 15: atomicity
- **Observed:** Serena rename/file-move/inline each ran as one refactor operation after symbol selection.
- **Built-ins:** a single `apply_patch` can update multiple files atomically as a patch, but it **cannot discover missed sites**; semantic completeness remains manual.
- **Delta:** Serena's advantage was not transactional all-or-none patching; it was **scope computation**.
**Verdict:** Serena improved semantic completeness more than patch atomicity.
#### Task 16: success signals
- **Observed Serena success outputs:** `OK` for body replacement, `"Success"` for rename, JSON result for move, `{"status":"SUCCESS"}` for inline.
- **Observed built-in success outputs:** only indirect evidence via `git diff` / clean revert.
**Verdict:** Serena gave clearer machine-readable success signals for refactors than the built-ins did.
- **Precision of matching:** Serena's reference search answered "who uses this in code?" better than `rg`, which mixed real uses with prose/comment matches.
- **Scope disambiguation:** Serena targeted exact symbols (`AutoLauncher/toggleAutoLaunch`, `InferenceBackend`) rather than relying on unique text strings.
- **Atomicity:** Serena computed and updated semantic scope in one refactor call; built-ins could batch edits, but only after manual scope discovery.
- **Semantic queries vs text search:** hierarchy and references were the strongest examples. Built-ins could reconstruct them, but only with manual interpretation.
- **External dependencies:** after indexing was available, Serena resolved desktop TypeScript dependencies into package-local declaration files under `desktop/node_modules` and Rust dependencies into external Cargo sources such as `url` and `zeroize`. Built-ins could still reach those files, but only after manual package-root or registry-path discovery.
- **Monorepo effect:** this repo magnified Serena's dependency-lookup value because "the dependency source" was not at one obvious global root. Serena jumped from app code to the right package-local or registry-backed dependency context directly.
**Verdict:** Serena improved correctness by narrowing work to exact symbols and by resolving dependencies across monorepo boundaries.
## 6. Workflow effects across a session
- **Advantages compounded** when I stayed in symbol space. Example: `get_symbols_overview(main.ts)` produced symbol names that I later reused for `find_symbol(createMainWindow)`, `rename(openStreetMapUserAgent)`, and `inline(waitForRendererDevServer)`.
- **Built-in workflows required refreshes**. Across repeated `main.ts` experiments, I repeatedly had to reacquire ranges with `view`/`rg` before editing because prior line-based context was no longer trustworthy.
- **In the monorepo, Serena also compounded by removing package-boundary bookkeeping.** In the desktop app I could jump from `main.ts` into Electron and `next-electron-server` declarations without first reasoning about workspace roots; in `rust/core` I could jump into Cargo-registry dependencies through external symbol handles instead of manually reconstructing registry paths from `Cargo.lock`.
- **The compounding effect disappeared** for tiny edits and non-code work, where built-ins were already direct and minimal.
- **One tradeoff compounded too:** some Serena refactors carried formatting side effects (notably `inline`), so the semantic benefit does not guarantee a surgically small diff.
**Verdict:** Serena's advantages compound most in code-centric monorepo sessions, where symbol reuse and dependency jumps save both re-reading and package-root discovery work.
## 7. Unique capabilities
| Capability with no practical one-step built-in equivalent | Frequency | Impact |
| --- | --- | --- |
| **Semantic cross-file rename from a single symbol definition** | Medium | High |
- Reading non-code files like `desktop/package.json`
- Free-text search such as `ente://app` or URL strings
- Git inspection / diff / cleanup
- Config/package/changelog/docs/notebook reading
- Exact textual patching once the line range is already known
In this session, these built-in-only tasks were **roughly 40% of the total operational steps by count**, but they were usually the low-complexity steps around the more valuable semantic work.
**Verdict:** A substantial share of everyday terminal work remains built-in-only, but Serena targets the higher-value symbol-heavy slice rather than the whole session.
## 9. Practical usage rule
Use **Serena first** when the task is about a **code symbol** and especially when it spans **multiple files, references, or a whole symbol body**. Use **built-ins first** when the task is about **text, config/docs, free-text search, git state, or a 1-3 line local tweak**. The highest-yield mixed workflow in this repo was: **discover/refactor with Serena, inspect non-code and do tiny patches with built-ins**.
**Verdict:** Choose Serena for symbol semantics and built-ins for text locality.
Reference in new issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.