Skip to content
PROJECT 02 / 11

Turboscore

Used-car assistant for Norway and Sweden. Plain-language search, clarifying questions, and a 1–100 score built from published components.

→ Every car scored 1–100 on twelve published components.

  • LLM
  • Search
  • Evals
  • Public product

Visit live site


Turboscore helps people in Norway and Sweden find a used car, live at turboscore.nobuns.com. It collects listings from the largest listing portals in both countries, scores every car from 1 to 100, and lets you search the way you would ask a friend who knows cars. The interface is in Norwegian, Swedish and English.

The project’s first design rule is that Turboscore is an assistant that helps you find a car, not a search engine that has to return something. When a question is too broad to answer well, it asks a clarifying question. When it cannot filter on what you asked for, such as a specific piece of equipment the listings do not record reliably, it says so instead of pretending. Clarifying and advising count as answers, not failures.

That rule is enforced. Search quality is gated by a baseline test suite, and the project explicitly never grades search on “share of queries that returned results” or on speed. A known gap is recorded as a known gap, and the test is never loosened to make it pass. The same honesty applies to filters: a “blind spot monitor” filter that had no data behind it was removed rather than left to return plausible-looking results.

Rules first, a local model when needed

Most queries never touch a language model. A pattern extractor reads brand, model, fuel, budget, year and mileage straight from the text. Only when it finds nothing usable, or the query leans on a cultural reference (“a car like Elvis would drive”), does Turboscore ask a local LLM, running through Ollama on the same server, to extract search parameters. The model can add parameters but cannot override what the patterns found. It gets two attempts with short timeouts, and if both fail the search continues on the patterns alone.

┌─[ 00 QUESTION ]──────────────────────────────────────┐
│ > family car, automatic, under 300k                  │
└─────────────────────────┬────────────────────────────┘
                          ▼
┌─[ 01 UNDERSTAND ]────────────────────────────────────┐
│                                                      │
│   patterns ──> local LLM only if patterns find none  │
│                                                      │
└─────────────────────────┬────────────────────────────┘
                          ▼
┌─[ 02 DECIDE ]────────────────────────────────────────┐
│                                                      │
│   too broad? ──> ask · can't filter? ──> say so      │
│                                                      │
└─────────────────────────┬────────────────────────────┘
                          ▼
┌─[ 03 ANSWER ]────────────────────────────────────────┐
│                                                      │
│   matching cars, ranked by TurboScore 1–100          │
│                                                      │
└──────────────────────────────────────────────────────┘
Patterns first, a local model as a fallback, and a clarifying question when that is the better answer.

A test that can fail

The baseline suite holds 20 queries, each with the intent and the search parameters it is expected to produce. For six weeks in mid-2026 it reported green while asserting nothing: the thresholds lived outside the test runner, so it passed at any accuracy. The fix in August was to make it assert 100 % intent accuracy and the response-time budget, and to break an expectation on purpose first to see a red build. A test that has never been seen to fail is not yet a test.

Its first real catch was the local model inventing car brands. Asked what to buy for towing a small boat, a question with nothing to filter on, the model answered with seven American makes from its own priors, narrowing 216 000 cars to 180 and presenting them as the answer. The brand validator that should have caught this accepted any brand that existed in the database, which is every brand. A brand must now be named in the query to be trusted, with one exception list for cultural references, where recalling an unnamed brand (James Bond, Aston Martin) is the whole job. The same list decides whether to call the model at all. Other invented parameters on advisory questions are recorded as a known issue in the suite, not hidden by loosening it.

The score, and making its claims true

TurboScore is a weighted score over twelve components: mileage, age, registered usage history, brand reliability, price against the market, fuel efficiency, safety, maintenance cost, listing condition, time on the market, recent price drops, and the EU roadworthiness inspection (EU-kontroll) from Statens vegvesen. The weights vary by the kind of car, so a classic is judged differently from a daily driver. When a component has no data for a car, its weight drops out and the rest are rescaled, so a Swedish car without a Norwegian inspection record is neither rewarded nor punished for a gap in our data.

In July 2026 an audit found that the landing page promised more than the code delivered: EU-kontroll data, insurance and tax costs, and a value forecast. Two of those were made true. The Vegvesen integration had been reading the wrong field names and silently returning nothing; it was fixed and now feeds the score from an hourly job. The rest were removed from the page. Two components that were defined in the configuration had never affected a single score, and now do. The methodology page now publishes the weights the scorer actually uses.

Two countries, one comparison

Prices arrive in NOK or SEK and Swedish mileage in mil (10 km), so every comparison runs on normalised values, never on the raw listing fields. Scheduled scrapers keep listings fresh, detect sold cars, and feed the price and time-on-market signals.

Anyone can use the chat assistant without an account, up to a limit per visitor. Public pages are prerendered for search engines, fonts are self-hosted, and each build ships its own strict Content Security Policy.

Stack

  • Flask API, MariaDB, Redis; scheduled scrapers instead of a task queue
  • React frontend, prerendered public pages, Norwegian, Swedish and English
  • Ollama running a small local model on the production server
  • Statens vegvesen vehicle data for the EU-kontroll component
  • Self-hosted on Linux, deployed on push

Status: public product, built since August 2025. Spotscore shares its architecture.

← back to work

esc
↑↓ nav ↵ open esc close