Give LlamaCppProvider an inspect() so local fingerprints use /props - #2033
HarshRajSinghania wants to merge 2 commits into
Conversation
Selection already calls provider.inspect() for local models, but LlamaCppProvider had no such method. The AttributeError was swallowed and fingerprints fell back to the filename. inspect() reads /props once at selection time and hands the payload to from_llamacpp(), which prefers general.parameter_count and only then the name. It does not share the capabilities cache used by chat. Fixes MODSetter#1989
|
@HarshRajSinghania is attempting to deploy a commit to the Rohan Verma's projects Team on Vercel. A member of the Team first needs to authorize it. |
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Closing as superseded by #2034, which adds the same #2034 won on two points:
One other thing worth mentioning for future PRs: the diff escapes the quotes in an existing docstring ( The |
Summary
_collect()already callsprovider.inspect(model_name)for a local model.LlamaCppProviderhad noinspect(), so theAttributeErrorwas caught and the fingerprint came from the filename only.This adds
inspect()on the adapter. It reads/propsonce at selection time and passes the payload to the existingfrom_llamacpp(), which prefersgeneral.parameter_countand falls back to the filename only when the runtime states none.Motivation
Fixes #1989. A local 70B renamed without a size in the filename was being tiered as
compact.Implementation
LlamaCppProvider.inspect()callsRouterClient.props()andfrom_llamacpp(). It does not write the capabilities cache; that cache exists for per-turn chat, and this runs once when a model is chosen.FakeRoutercan state aparameter_counton/propsso the unit tests can drive a name with no size against a runtime that does state one.The
try/exceptsafety net in_collect()is left in place so a failed/propsread still falls back tofrom_name().Testing
The new unit tests assert:
renamed-weightswithgeneral.parameter_count = 70e9fingerprints as 70.0Binspect()andcapabilities()each make their own/propscallI did not run the full
uv run pytest -m unitsuite in this environment (no project venv / llama.cpp test extras). The added tests follow the existing FakeRouter style intest_provider.py.Please run from the issue:
High-level PR Summary
This PR fixes a fingerprinting issue where local LLM models without size information in their filename were being incorrectly tiered. It adds an
inspect()method toLlamaCppProviderthat reads parameter count from the llama.cpp runtime's/propsendpoint, ensuring accurate model classification based on actual model properties rather than just filename parsing.⏱️ Estimated Review Time: 5-15 minutes
💡 Review Order Suggestion
surfsense_local/backend/tests/unit/llm/providers/llamacpp/fake_router.pysurfsense_local/backend/tests/unit/llm/providers/llamacpp/test_provider.pysurfsense_local/backend/modules/llm/providers/llamacpp/provider.py