TypeScriptLanguageServer._find_representative_source_file scanned files directly
adjacent to tsconfig.json before checking a src/ subdirectory, so a same-level tool
config that the tsconfig excludes (vitest.config.ts, jest.config.ts, etc.) could be
picked as the file used to trigger project activation for an additional workspace
folder. The wrong inferred TypeScript project then loads, and cross-package
find_referencing_symbols queries silently return {} instead of raising or warning.
Reorder the scan to walk the src/ subtree first (recursively, respecting the usual
ignored-directory rules) and only fall back to a same-level file when no src/
directory exists, matching the fallback the reporter's issue proposed.
Fixes#2090
Signed-off-by: Amir Fathi <amirfathi.me@gmail.com>
* fix(csharp): fold newly created files into the loaded Roslyn project
Roslyn only learns which files belong to a project from the
solution/open and project/open notifications _open_solution_and_projects
sends once at startup. The generic didChangeWatchedFiles/open-close cycle
that poll_and_notify sends for every backend on file creation does not
make it re-evaluate the project, so a .cs file created after the project
was already indexed stayed a standalone Miscellaneous Files document,
producing phantom diagnostics (e.g. an incorrect "using directive is
unnecessary") instead of the real ones.
Add an overridable SolidLanguageServer.notify_files_created hook, called
by LanguageServerFileChangeNotifier.poll_and_notify before it opens newly
created files. CSharpLanguageServer overrides it to resend the same
solution/project notifications and wait for Roslyn's own reload-complete
log line before the file is opened; every other backend keeps its current
behavior via the no-op default.
Fixes#1961
* fix(test): use get_language_server_manager_or_raise in csharp fold test
project.language_server_manager is typed LanguageServerManager | None, so
calling .iter_language_servers() on it directly failed ty's unresolved-attribute
check. Agent already exposes a non-Optional accessor for the same value.
_wait_for_cross_file_references_if_needed latched after the first cross-file
query and never waited again, even when a later query opened a file from a
project tsserver had not loaded yet and started a fresh $/progress cycle.
find_referencing_symbols and find_references would then return whatever
tsserver had indexed so far, silently partial, with nothing distinguishing it
from a complete answer.
The latch still covers the first query's start-grace wait unchanged. A later
query now also checks whether a progress token is currently active and, if
so, waits for it to drain before returning.
Fixes#1937
search_for_pattern resolved each match's line number by rescanning the
file from the beginning (TextUtils.get_line_from_index walks a
TextStepper from index 0), making coordinate resolution O(n) per match
and O(n*m) for m matches. search_text now precomputes line start
offsets once per file (O(n)) and resolves each match via binary search
(O(log n) per match).
The precomputed table mirrors TextStepper line semantics exactly:
\n, \r\n and bare \r are line separators, and an index pointing at the
\n of a \r\n pair resolves to (next line, 0) — the same rule
TextUtils.get_line_col_from_index implements, now pinned by regression
tests (including the mixed line-ending edge cases from the
insert_text_at_position fix).
On a synthetic 12k-line file (494 KB) with 723 matches, match-coordinate
resolution drops from ~1764 ms to ~2 ms (~1000x); the one-off
precompute costs ~1.1 ms. Small files regress nothing (100 lines, 5
matches: 70 µs -> 11 µs). Results are byte-for-byte identical to the
previous implementation.
Note: compatible with PR #1899 (interruptible agent regexes): that PR
touches the re.compile call site in search_text and the ContentReplacer
helpers, this PR touches the post-match line-coordinate resolution; the
two changes are complementary and rebase cleanly onto each other.
* perf(search_text): add reproducible benchmark script
scripts/profile_search_text.py regenerates the synthetic benchmark
content with a fixed seed (no real-world code), asserts both resolution
paths produce identical results, and prints the before/after timings
quoted in the PR description. Default: 12k lines / ~1830 matches;
optional CLI args scale the file down for quick runs.
* refactor(ls): encapsulate cached text coordinates in TextCoordinates abstraction
Address review feedback on the search_text line-lookup PR: the previously
added module-level helpers duplicated TextStepper's line-ending logic and
exposed raw offsets as a low-level data structure shared between two
functions, violating the project's software design principles (each
concern in exactly one home; dataclasses instead of tuples for simple
data storage).
Introduce in solidlsp/ls_utils.py:
- LineCol: frozen, kw-only dataclass for a 0-based line/column position
- TextCoordinates: encapsulates the cached table of line start offsets,
built via TextStepper (step_line + line_start_idx), with a single
public method line_col_at_index() resolving an index via binary
search, including the CRLF rule (an index pointing at the '\n' of a
'\r\n' pair denotes the beginning of the next line, column 0) and
InvalidTextLocationError for out-of-range indices
search_text now instantiates TextCoordinates once per file and resolves
each match's coordinates through the public method; the private helpers
(_compute_line_starts, _line_col_at_index) are removed.
Equivalence with TextUtils.get_line_col_from_index verified across
mixed line endings, empty content and boundary indices (0 mismatches).
Tests are behaviour-anchored: the implementation-coupled parametrized
test importing the private helpers is replaced by public-contract tests
of TextCoordinates (test/solidlsp/test_ls_utils.py) and absolute line
assertions through the public search_text API. The benchmark script
uses the public API as well.
Benchmark (synthetic 12k lines, ~723 matches, best of 5): coordinate
resolution ~1764 ms -> ~2 ms; precompute ~1.1 ms; small files unchanged.
Semantics are unchanged; results are identical to the previously
proposed implementation and to the upstream reference implementation.
* fix(test): rename test_ls_utils to test_text_coordinates to avoid basename collision
The new test file collided with the pre-existing upstream
test/solidlsp/util/test_ls_utils.py: pytest (prepend import mode, no
__init__.py in the test dirs) requires unique basenames, and collecting
both modules produced an import-file-mismatch error that aborted
collection with exit code 2 across every CI matrix job.
Rename to test_text_coordinates.py (unique, and matches the class under
test). Full-suite collection verified locally (2252 tests, 0 errors).
* Bridge external language server adapters through the registry
The plugin mechanism (ExternalLanguageServerId + LanguageServerRegistry
via solidlsp.language_server_registration entry points) was added but
three places that consume language IDs still only iterate the built-in
LanguageServerId enum, making externally-registered adapters
unreachable from project.yml and the CLI.
This change:
- _determine_project_language_servers: also scans externally-
registered LSes for auto-detection.
- serena project create --ls: accepts registry keys as fallback.
- No behavior change for built-in language servers.
Motivation: enables first-class support for language servers like
vhdl_ls that ship as separate packages rather than being vendored
into the main repo.
* Widen signatures for external language server support
Follow-up to fix/external-ls-registry:
- compute_language_server_support_composition: accept and return
LanguageServerIdLike instead of LanguageServerId enum members only,
since externally-registered adapters are also LanguageServerIdLike.
- ProjectConfig.autogenerate: same widening for the languages parameter.
- Two string-formatting sites in _determine_project_language_servers
use .get_key() instead of .value so they work for either type.
Type checker (ty) clean on src/serena and src/solidlsp; pre-existing
diagnostics unrelated to this PR remain.
* Refine based on review: registry as SoT, widen get_ls_priority
Following @opcode81's review feedback:
- Add LanguageServerRegistry.iter_registered_ls_ids() — a single
iterator that yields all registered LSes (built-in enum + externals).
- _determine_project_language_servers uses the new iterator instead
of the split 'enum + dedup-externals' loop.
- get_ls_priority now accepts LanguageServerIdLike, so user-configured
priorities via serena_config.ls_priorities apply to external LSes
too. Internally uses ls.get_key() (defined on both enum and
ExternalLanguageServerId).
- project create --ls drops the try/except fallback: resolves via
registry directly. Unknown keys now report the full registry key
list in the error message.
Net diff vs previous attempt: +21/-27 vs +33/-14 (simpler).
* Bridge external language server adapters through the registry
The Dart analysis server treats rootUri as an additional analysis root on top of
workspaceFolders, with no de-duplication, so on a monorepo root the whole tree is
analysed and the server burns CPU at idle. Always omit rootUri/rootPath and rely on
workspaceFolders alone.
Rename propagation
* `rename_memory_and_propagate_references` enumerated `get_full_list()`, which
includes the memories matched by `read_only_memory_patterns`.
* Writing one of those raises `PermissionError` in `_check_write_access`, so a
rename of a memory that was referenced from a read-only memory failed *after*
`move_memory` had already applied the rename.
* The memory graph was then half-updated: the renamed memory existed under its
new name while the writable referrers still held `mem:OLD_NAME`, and retrying
the rename failed with "Memory not found".
Fix
* In a tool context, propagate only into the memories that accept writes, sorted
to keep the enumeration order that `get_full_list()` provided.
* Outside a tool context nothing changes, so read-only memories are still updated
when the user renames through the CLI.
* A reference which thereby remains in a read-only memory is still reported as
stale by `validate_referential_integrity`, so it is not hidden.
Documentation
* `docs/02-usage/045_memories.md` promised that the tool rewrites every reference across all
memories. That is now only true for the memories the agent may write, so the caveat is stated
where the promise is made.
Tests
* Add a regression pair: the tool-context rename completes and rewrites both
occurrences in a writable referrer, while the CLI context additionally rewrites
the read-only one.
* fix(tools): expand unmatched-group $!N backreferences to the empty string
A $!N backreference in a regex-mode replacement template refers to the Nth
group of the search expression. When the group exists but did not participate
in the match (e.g. it sits inside an optional construct that was skipped), the
expansion emitted the literal template text instead of the empty string.
Observed in practice when an agent replaced 15 occurrences in an MQL header
with a template containing EA_INPUT$!1(...): the group was optional and never
participated, so every site came out containing the literal EA_INPUT$!1(...).
A reference to a group the expression does not define at all now raises a
clear ValueError instead of a raw IndexError - which also crashed
literal-mode replacements whose template contained $!N, since literal mode
compiles an escaped pattern without any groups.
* chore(ci): re-trigger test workflow
The previous run failed in native (macos-latest) on
test_cpp_basic.py::test_get_document_symbols[cpp_ccls] ("Expected 'main'
in document symbols, got: []"), a ccls language-server startup/indexing
flake unrelated to this PR's changes (text_utils.py and its tests only).
* fix(tools): pass literal-mode replacement through verbatim (no $!N expansion)
ContentReplacer ran the $!N backreference expansion on the replacement
template in both modes, so a literal-mode replacement whose template
contained a $!N sequence (e.g. documenting the convention itself in a
memory) failed with a backreference error instead of writing the text.
This contradicted the tools' documented contract ('the replacement string
(verbatim)' in literal mode) and the behavior of the dry-run path, where
MultiFileContentReplacer.find_occurrences already gated the expansion on
regex mode.
The expansion is now gated on the mode, mirroring the dry-run path; the
ambiguity validation still applies in both modes. The literal-mode test
that asserted the crash is inverted accordingly and the changelog entry
is corrected.
SerenaConfig.project_names / project_paths were cached_property values that were
never invalidated after projects were added or removed mid-session, so
user-facing project lists and error messages stayed stale. They are cheap to
derive, so drop the caching (plain properties) instead of invalidating.
CSharpLanguageServer._open_solution_and_projects scanned the whole
repository root and opened every .csproj it found, without consulting
the project's ignore settings.
On repositories that vendor third-party or sample C# projects this
opens projects the language server cannot restore. The cost is paid on
every server start, and the resulting restore failures bury the
diagnostics of the projects the user actually works on. Measured on an
Unreal Engine source tree: 245 projects opened, 53 of them under
Engine/Source/ThirdParty, which Roslyn cannot build; each restart
emitted thousands of NuGet advisory lines and ended in
`The "Csc" task could not be initialized`.
SolidLanguageServer.is_ignored_path already implements exactly this
check, and CSharpLanguageServer already overrides is_ignored_dirname,
so the ignore settings were being honoured everywhere except here.
This applies the existing check at project discovery.
ignore_unsupported_files=False is required because a .csproj is not
itself a C# source file, and would otherwise be excluded on file type
rather than by the ignore patterns.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Dr. Dominik Jain <dominik.jain@oraios-ai.de>
* fix(csharp): stop tuple-typed properties losing their name to the method branch
_extract_base_name_and_type split Roslyn's "Name : Type" property names by
checking for a literal '(' anywhere in the raw string. A C# tuple type is
written with parentheses ("(int X, string Y)"), so a tuple-typed property
tripped that guard and fell into the method branch instead, which kept the
trailing " :" as part of the reported name (e.g. "Position :"). Since
find_symbol defaults to exact name-path matching, such a property becomes
unfindable by its real name.
The guard now only checks for '(' in the identifier segment before the
first " : ", not the whole string, so a parenthesis inside the type
annotation no longer affects branch selection.
* fix(csharp): bump the high-level symbol cache fingerprint
_normalize_symbol_name's output changed in this PR, and
_document_symbols_cache_fingerprint gates the cache that stores its
result; leaving the version at 1 would keep serving the old, corrupted
names to anyone with an existing on-disk cache.
The previous default, 262.9593.0, is a JetBrains EAP-style build that has
expired: intellij-server exits at startup with "This build of intellij-server
has expired", so every Kotlin symbolic tool fails. 263.4702.0 is the newest
release on Kotlin/kotlin-lsp and starts normally.
Uses the same archive layout and CDN path as 262.9593.0. Hashes regenerated
with scripts/update_downloaded_dependency_hashes.py.
Fixes#2008
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Project.gather_source_files walks the tree with os.walk, which hands back
directories and files separately, and then calls is_ignored_path on each
one. That method re-derived file-ness from the filesystem: an
os.path.exists and an os.path.isfile in _is_ignored_relative_path, plus
an os.path.isdir in match_path. Three syscalls per path, for an answer
the caller already had.
is_ignored_path, _is_ignored_relative_path and match_path now accept an
optional hint, and gather_source_files supplies it from the os.walk
split. The parameter defaults to None, which determines file-ness from
the filesystem exactly as before, so no existing caller changes
behaviour.
Measured on a repository with 97,549 tracked source files out of 707,889
total (Unreal Engine source plus three game projects; Windows 11,
Python 3.13):
gather_source_files() 67.8s -> 11.6s
The returned file list is byte-identical before and after (sha256 over
the sorted relative paths, 10,390,077 bytes).
For context on where the time went: a bare os.walk of the whole 708k-file
tree takes 10.7s, so this was never I/O-bound. cProfile over 15,000 real
is_ignored_path calls attributed 3.41s to nt._path_exists,
nt._path_isfile and nt._path_isdir - about 82% of the per-call cost.
The new test asserts equivalence rather than specific verdicts: for every
path in a fixture tree, across several ignore configurations, the hinted
call must agree with the unhinted one. It covers the two cases where
guessing file-ness from the name would go wrong - a directory with a
suffix, and an extensionless file - and was checked against a deliberate
inversion of the hint, which it catches.
Refs #2077
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Dominik Jain <dominik.jain@oraios-ai.de>
Previously, Tool.apply_ex derived a session ID from the MCP session object (or
"global" with no context) and injected it into any apply() method that declared
a session_id parameter. This is being changed because:
* the new MCP SDK v2 no longer provides session identifiers
* handling it internally is more robust anyway, since clients did not
consistently use sessions
Session handling
- Remove implicit session ID injection from Tool.apply_ex, along with the
supporting _is_session_aware property and SESSION_ID_PARAM_NAME skip logic
- SerenaAgent.create_system_prompt() now creates the session itself and reports
its id, instead of receiving session_id as an argument
Tool signatures
- InitialInstructionsTool.apply() and ActivateProjectTool.apply() now declare
session_id explicitly and rely on the LLM to pass it, rather than having it
injected
- Rename SerenaReplTool.apply()'s session parameter to session_id for consistency
with the other tools
Resolves#2061
* Many case differentiations in the agent code were replaced by
method calls in the newly introduced LanguageBackend abstraction
* The registry allows new backends to be added dynamically
(via Python packages that implement a specific entrypoint)
- Require a valid Bearer token on heartbeat and project-query requests.
- Send the configured auth_secret from clients and the query-project tool.
- Test accepted and rejected credentials, including real HTTP client requests.
- Set configuration permissions to 0600 before reading and log adjustments.
- Log chmod failures and continue loading; leave Windows permissions unchanged.
- Generate a random UUID when auth_secret is missing, null, or empty.
- Persist generated secrets and preserve configured values across reloads.
- Document the shared secret and test generation, persistence, and direct construction.
- Rename the read-only context to read_project_context and add writable project_context.
- Dispatch LSP-backed edits remotely and enforce read-only access in the calling context.
- Allow project-server facade calls independently of the target project's API restrictions.
- Update dispatch coverage for both backends and access modes; remove obsolete rejection assertions.
SerenaDashboardTrayManager._update_menu() called pystray's Icon.update_menu()
on whatever thread reached it. Four call sites are off the main thread: the
Flask handlers for /register, /update_project and /unregister, and
_alive_check_loop.
pystray does no marshalling. Icon.update_menu() calls the backend directly and
pystray/_darwin.py goes straight to NSStatusItem.setMenu_(), which AppKit
requires on the main thread. On macOS versions that enforce it the tray manager
traps with SIGTRAP inside the request handler, so the agent logs "Failed to
register with tray manager: Remote end closed connection without response" and
the tray icon never becomes usable.
Menu refreshes now go through PyObjCTools.AppHelper.callAfter on Darwin, which
is asynchronous so no Flask handler blocks on the main run loop. Other
platforms call through unchanged. The helper is separate from _update_menu
because _open_dashboard and _run_viewer run on the same Flask threads and will
need it too.
No new dependency: PyObjCTools comes from pyobjc-core, already required on
macOS via pystray -> pyobjc-framework-Quartz -> pyobjc-core.
The crash itself could not be reproduced locally (macOS 26.6.2; the reporter is
on 27.0, everything else matching). Verified instead that the AppKit call moves
from a Flask worker thread to the main thread: NSThread.isMainThread() across
the three HTTP routes reads [False, False] before and [True, True, True] after.
A language server's cache directory was determined by the language_id rather than
the language server identifier's key. The two identifiers coincided in most cases.
Facade availability:
* facades can be optional (`Facade.from_api(..., is_optional=True)`), mirroring optional tools
* `Facade.is_enabled()` is derived: a facade is available iff it has at least one enabled method
* `ApiScope.is_facade_enabled` is thereby obsolete and removed
Opt-in rule (uniform for optional facades, excluded facades and optional methods):
* methods of a facade which is not included and optional methods require explicit inclusion
* all other methods are enabled unless explicitly excluded
Application:
* the `ext` facade is optional; the query-projects mode includes it (`included_apis: [ext]`)
Godot's GDScript parser can report a symbol's end column one column past the
line-end convention every other language server follows (closing a node's range from
the next lookahead token instead of the last consumed one, when that lookahead is a
synthesized newline); `replace_symbol_body` on the last function in a file silently
consumed the separating blank line as a result. `GodotLanguageServer` now corrects this
specific, measured overshoot when building its high-level document symbols (#1974)
---------
Signed-off-by: Amir Fathi <amirfathi.me@gmail.com>
The project registry in serena_config.yml can be added to from the CLI
('project create', 'project index') and from the MCP side
('activate_project'), but it could only be removed from on the MCP side,
via the 'remove_project' tool. From a terminal, hand-editing the file was
the only way to drop an entry.
'serena project remove <PROJECT_NAME|PROJECT_PATH>' takes the same
argument type as 'project index', so a project can be addressed by name
or by path. Only the registry entry is removed; the project's own files,
including its project configuration, are left on disk.
Adds SerenaConfig.remove_registered_project, which mirrors
add_registered_project and removes the given entry rather than the first
entry whose name matches. Resolving a path to an entry and then deleting
by name would remove the wrong project when two registered directories
carry the same project_name.
Refs #2029
`serena.analytics` imported `anthropic.types` at module level for two type
annotations, so every `import serena.cli` / `import serena.agent` paid for
loading the whole anthropic package (2-4 s on the reporter's Windows machine)
even with the default CHAR_COUNT token estimator. That delay can push MCP
startup past a client's initial tool-discovery window.
The import now lives behind `TYPE_CHECKING`; `AnthropicTokenCount` already
imported `anthropic` lazily in `__init__`, and the request body is a plain
dict (what `MessageParam` produces at runtime anyway).
Tests: importing serena.cli / serena.agent / serena.analytics in a fresh
interpreter leaves `anthropic` out of `sys.modules`; the Anthropic estimator
still sends the expected count_tokens request.
Fixes#2012