Methodology · Forkelo Score v1
How a score is calculated.
Forkelo publishes one number for one dish at one restaurant. This is the whole of it — the evidence it is made from, the arithmetic that turns that evidence into a number between 0 and 10, every constant with the reasoning behind it, and three real calculations traced row by row.
Abstract
A restaurant rating is an average over everything a restaurant does. A dish rating should be an average over one dish. Forkelo builds the second from primary evidence: individual, dated, dish-specific mentions, each classified for sentiment, for how strongly it recommends, and for how sure the reader was that it referred to this dish at this restaurant. Each mention is scored on a 0–1 valence scale and weighted by recency and reader confidence. Their weighted mean is shrunk toward a neutral prior in proportion to how little evidence there is, given a small bonus when independent kinds of source agree, and printed on a 0–10 scale to one decimal place.
How certain that number is, is never mixed into it. Confidence is computed separately from the total weight of evidence and the number of independent source types, and printed beside the score. The whole calculation is a pure function of the evidence rows: no restaurant star ratings, no popularity, no personalisation, no editorial thumb, and nothing that can be bought.
Why a dish needs its own number
Restaurant ratings answer a question nobody actually has. “How good is this restaurant?” is an average over the food, the service, the wait, the room, the parking and the day somebody's order was wrong. It is a single number standing in for dozens of independent things, and the moment you want to know which of four taquerias makes the best carne asada burrito, it is close to useless. A 4.6-star restaurant can serve a forgettable burrito. A 3.9-star restaurant with slow service and no parking can serve the best one in the county.
The failure is not that star ratings are inaccurate. It is that they answer at the wrong level of aggregation. Averaging over a restaurant destroys exactly the variation you are shopping on — the difference between its dishes. No amount of arithmetic applied afterwards can recover it, because the information was thrown away before the number was formed.
The rule the whole system is built around: a restaurant's overall rating is never evidence about a dish, and cannot become an input to a Forkelo Score by any path. Not scaled, not blended, not used as a tie-break, not used as a prior. The only inputs are mentions of the specific dish.
Everything below follows from four design constraints, in this order of priority:
- Traceable. Every score decomposes into the rows that produced it. A number nobody can take apart is a number nobody should trust.
- Conservative under thin evidence. Two enthusiastic mentions must not produce a 9.8. The cost of a wrong #1 is much higher than the cost of a good dish sitting at 7.6 until it is properly researched.
- Explainable in one sentence per step. Nothing here is learned, fitted, or tuned against engagement. Every constant is a stated judgement that can be argued with — including on this page.
- Uncorruptible by volume. Loudness is not quality. A dish cannot climb by being mentioned more in the same place, only by being mentioned well, recently, and in more than one kind of place.
What an aggregate score has to decide
The reference point most people arrive with is the Tomatometer: one number, assembled from many independent critics, that you can check in three seconds and then argue about for an hour. Films have had a number like that for a generation. Restaurants have had one for rather longer. The individual dish — the thing people actually choose, and the thing they will cross town for — has never had one, and that gap is what Forkelo is built to fill.
Building such a number means making three choices. None has an obvious answer, and a film-style consensus score answers all three differently from Forkelo. What follows compares the shape of the two designs, not anyone else's implementation, whose thresholds change and are not ours to describe.
One: count the voices, or weigh them?
A consensus score is a headcount. Each critic casts one vote, positive or negative, and the published number is the share that came out positive. The virtues are real: it cannot be misread, one impassioned voice cannot swing it, and it gets steadier as the sample grows.
What a headcount throws away is intensity. “The best film of the decade” and “worth a matinee” are the same vote. In food that loss is fatal, because “the best asada burrito in San Jose” is not a louder version of “the burrito was fine” — it is the specific claim a ranking exists to test. So Forkelo weighs rather than counts: a superlative outweighs a compliment (section 6), a fresh source outweighs an old one (section 7), and a clear reading outweighs an ambiguous one.
Forkelo does publish the headcount. It just is not the score. Every ranking card carries a line like “12 dish-specific mentions · 83% positive · 3 source types”, and that percentage is exactly the consensus measure: the share of mentions that were positive. It runs beside the score as context, because on its own it cannot separate two dishes no diner would confuse.
| Dish | Mentions | % positive | Forkelo Score |
|---|---|---|---|
| Example A | 2 | 100% | 7.3 · Early ranking |
| Example C | 30 | 100% | 9.6 · Early ranking |
| Example B | 10 | 80% | 8.5 · Medium confidence |
A and C are both unanimous, and a percentage says they are the same dish. They are not remotely the same dish. Meanwhile B, the only one of the three carrying a bad review, outranks the unanimous one — which is the right answer, and not one a headcount can produce.
Two: does time count?
A film is finished. The print does not change, so a review written in 1998 is as true of it now as it was then, and a consensus score is right to keep counting that review forever.
A kitchen is never finished. It changes chefs, suppliers, owners and standards, and a rave about a dish is a claim about a kitchen on one particular Tuesday. This is the largest structural difference between rating a film and rating a plate of food. It is why every piece of evidence here decays (section 7), and why a Forkelo Score moves even when nothing new has been filed (section 17).
Three: where does “enough evidence” live?
Every aggregate score eventually needs a way to say this number is thin. The film convention is a badge: clear a quality bar and a volume bar together and the work is marked; fall short of either and it is simply unmarked. A badge is legible at a glance, which is its whole appeal, but it files the failure somewhere nobody looks — a film with four reviews and a film with four hundred are both unmarked, and nothing on the page says which is which.
Forkelo prints the second number instead, always, beside the first. 9.4 · High confidence and 9.6 · Early ranking are both published, both sortable, and both mean something exact (section 12). The consequence is one a badge cannot produce: a dish can sit at the very top of a ranking and be marked as thinly evidenced in the same breath, on the same card, rather than by the quiet absence of a mark.
What Forkelo is not borrowing
The ambition is that shape — a number for the dish that anybody can check and everybody can argue with — and not its mechanics. A Forkelo Score is not a percentage, is not a headcount, certifies nothing, and cannot be bought (section 16). It is an estimate of how strongly, how recently and how widely one dish is recommended. The rest of this page is what that sentence costs in arithmetic.
The unit of observation
The atom of the system is an evidence row: one source, saying one thing, about one dish, at one restaurant, on one date. Nothing else enters the calculation. A dish with forty evidence rows has forty of these; a dish with none has no computed score at all.
A row is filed by a research agent that reads a source and records a structured observation about it. It never stores the source's text — the summary field is the researcher's own words, which is a licensing constraint before it is a style preference.
| Field | Values | Enters the score | Role |
|---|---|---|---|
| sentiment | positive · neutral · negative | Yes | Sets the row's base valence. |
| recommendation_strength | none · weak · moderate · strong | Yes | Raises the valence of a positive row. Not sentiment: it measures how emphatic the recommendation is. |
| confidence | low · medium · high | Yes | How sure the researcher is that this source means this dish at this restaurant. A property of the reading, not of the dish. |
| source_date | a date, or absent | Yes | Drives the recency weight. |
| source_type | review_site · reddit · blog · editorial · social · menu · forkelo_user · other | Yes | Independence. Distinct types are treated as independent corroboration. |
| attributes | short noun phrases | No | Counted by sentiment into the “what people mention” and “criticism” lists on a dish page. Never moves the number. |
| summary | ≤ 500 chars, researcher's words | No | The audit trail for a human reviewer. |
| source_url | a URL, or absent | No | Provenance, and the uniqueness key that stops one source being filed twice for one dish. |
What counts as a mention
The boundary is narrow on purpose. Evidence must be about the food, at that restaurant, specifically enough to be checked.
| Source text | Filed? | As |
|---|---|---|
| “The carne asada burrito here is incredible.” | Yes | positive, strength none |
| “Best asada burrito I've had in San Jose.” | Yes | positive, strength strong |
| “Burrito was fine, nothing special.” | Yes | neutral |
| “Way too much rice, barely any meat.” | Yes | negative |
| “Great restaurant, friendly staff.” | No | about the restaurant, not the dish |
| “4.4 stars, 1,200 reviews.” | Never | an aggregate, and the one thing that is forbidden |
Neutral and negative rows are filed as diligently as praise. A database of only compliments produces rankings nobody believes, and the criticism list is a feature of a dish page rather than an embarrassment. A researcher who files only the good mentions has not been generous; they have broken the denominator.
Notation
One dish has a set of evidence rows E = {r1 … rn}. For each row ri:
| Symbol | Name | Range | Meaning |
|---|---|---|---|
| vi | valence | 0.10 – 1.00 | What the row says about the dish. |
| ρi | recency weight | 0.20 – 1.00 | How much its age lets it count. |
| κi | reader-confidence weight | 0.5, 1.0, 1.2 | How much the researcher's certainty lets it count. |
| wi | row weight | 0.10 – 1.20 | ρi × κi. |
| W | evidence weight | ≥ 0 | Σ wi. The size of the evidence base, in units of “one fresh, medium-confidence mention”. |
| q | quality | 0.10 – 1.00 | The weighted mean valence. |
| ŝ | shrunk quality | 0.10 – 1.00 | q pulled toward the prior by an amount that depends on W. |
| S+ | agreeing sources | 0 – 8 | The set of distinct source types with at least one positive row. |
| a | agreement bonus | 0 – 0.08 | Corroboration across independent kinds of source. |
| S | the Forkelo Score | 0.0 – 10.0 | What the site prints. |
The calculation, end to end
Six lines, in order. Everything after this section is the justification for one of them.
In line 1 the bonus applies only to positive rows and v is capped at 1.0. Read the whole as one sentence: take what each mention says, weight it by how fresh and how trustworthy it is, average, pull that average toward the middle in proportion to how little evidence you have, add a little for independent corroboration, and print it out of ten.
The function is pure. Given the same rows and the same date it returns the same score on any machine, with no database read, no network call, no randomness and no state. That is what makes the worked examples in section 14 reproducible rather than illustrative.
Valence: what a mention is worth
Sentiment sets a base value on a 0–1 scale.
- 0.85positive
- 0.50neutral
- 0.10negative
Three decisions are embedded in those numbers.
Positive is 0.85, not 1.00. The top of the scale is reserved for a mention that is not merely positive but emphatic. If plain praise were worth 1.0, there would be nowhere left for “the best I've had in this city” to go, and the most useful distinction in food writing would collapse.
Negative is 0.10, not 0.00. Zero is an absorbing value: it says the dish is worthless, which no single disappointed diner has the standing to say. A pan is strong evidence, not proof, and 0.10 is enough to drag a score down hard — see section 15 — without letting one bad night annihilate a decade of praise.
Neutral is 0.50, which is below the prior of 0.55. This is deliberate and it surprises people. A genuinely unremarkable mention (“the burrito was fine”) is slightly worse news than no information at all, because somebody bothered to eat the dish and had nothing to report. Filing neutral evidence therefore very gently lowers a score. That is the correct direction.
The strength bonus
Recommendation strength is not sentiment. It measures how far a source is willing to stick its neck out, and it applies only to positive rows.
| Strength | Bonus | Valence | Sounds like |
|---|---|---|---|
| none | +0.00 | 0.85 | “The burrito is good.” |
| weak | +0.03 | 0.88 | “I'd get the burrito again.” |
| moderate | +0.08 | 0.93 | “Get the asada, trust me.” |
| strong | +0.15 | 1.00 | “Best asada burrito in the South Bay.” |
The bonuses are convex — 3, 5 then 7 hundredths — so each step up in emphasis is worth slightly more than the last. A superlative is not merely a bit more positive than a compliment; it is a different kind of claim, and it is also the claim most likely to be inflated, which is why the whole span from plain praise to superlative is only 0.15 wide. The gap between a good burrito and the best burrito in the city is worth about one-seventh of the gap between a good burrito and a bad one.
There are therefore exactly six values a row can take: 0.10, 0.50, 0.85, 0.88, 0.93, 1.00. Nothing else is representable, which is a feature — it keeps the classification job a small number of discrete judgements rather than an invitation to score a mention 7.5 out of 10 by feel.
There is no negative mirror of the strength bonus. “The worst burrito in San Jose” and “the burrito was bad” are both worth 0.10. Superlatives of disgust are cheaper to produce than superlatives of praise, and weighting them would make one furious diner the loudest voice on a page.
Weight: what a source is worth
Valence says what a row claims. Weight says how much that claim counts. The two are strictly separated: nothing about how much a source is trusted ever changes what it is understood to have said.
ρ ranges 0.20–1.00 by age; κ is 0.5, 1.0 or 1.2 by how sure the researcher was. A row can therefore weigh anywhere from 0.10 to 1.20 — a twelvefold span between the least and most useful piece of evidence in the system.
Reader confidence (κ)
This field is about the researcher, not the dish. A Reddit comment saying “the burrito there is great” posted under a thread discussing five taquerias is low: it is probably about one of them, and possibly this one. A named dish reviewed in a dated article is high.
- ×0.5low — might not be this dish, or this restaurant
- ×1.0medium — the default, and the unit of W
- ×1.2high — unambiguous
The scale is deliberately asymmetric. Doubt costs a row half its weight; certainty earns it only a fifth more. The reason is that the failure modes are not symmetric: over-crediting an ambiguous source puts a wrong claim on a public page, while under-crediting a clear one merely means more research is needed. A high row is worth 2.4 low rows.
Recency (ρ)
Restaurants change chefs, recipes, owners and standards. A rave from 2019 is a fact about a kitchen that may no longer exist. Forkelo decays evidence exponentially with an 18-month half-life, floored at 0.20.
months(d) is the age in months, computed as days ÷ 30.44 and clamped at zero, so a source dated in the future is treated as today rather than given extra weight.
| Age of source | ρ | In plain terms |
|---|---|---|
| today | 1.000 | full weight |
| 6 months | 0.794 | still nearly current |
| 1 year | 0.630 | worth about two-thirds of a fresh mention |
| 18 months | 0.500 | the half-life |
| 2 years | 0.397 | two 2024 mentions ≈ one 2026 mention |
| 3 years | 0.250 | a quarter |
| 3½ years+ | 0.200 | the floor; it does not decay further |
| undated | 0.700 | treated as about 9 months old |
Why 18 months. It is roughly the interval over which a kitchen can change enough to matter without the restaurant closing — a chef leaves, a supplier changes, portions shrink — and it makes the arithmetic legible to the people filing evidence: a 2026 mention counts about twice a 2024 one. A shorter half-life would make scores lurch as old evidence fell off a cliff; a longer one would let a restaurant coast for years on a reputation it no longer earns.
Why a floor at all. Pure exponential decay tends to zero, which would say that a 2015 review of a still-open restaurant contains literally no information. It contains some. Food changes; it does not vanish. The floor also prevents a pathology: without it, a dish whose evidence is uniformly ancient would have a total weight near zero, and shrinkage (section 9) would snap it exactly onto the prior — the system would confidently report 5.5 for every old dish, which looks like a measurement and is not.
Why undated evidence is worth 0.70. Undated evidence is not fresh and is not necessarily stale, so it takes a middling value — the equivalent of a source about nine months old. It is set slightly below the value of a genuinely recent mention so that failing to find a date carries a small, honest cost, without discarding otherwise good evidence.
Quality: the weighted mean
An ordinary weighted mean, with three properties worth naming because they are what stop the number from being gameable.
- It is a mean, not a sum. Adding more mentions of the same quality does not raise q at all. Volume only enters the score later, through shrinkage, and separately through confidence. Loudness alone buys nothing.
- It is scale-invariant. Doubling every row's weight leaves q unchanged. A dish with five high-confidence mentions and a dish with five medium-confidence mentions of identical content have the same quality — they differ in W, and therefore in how much of that quality survives shrinkage.
- It is bounded by its inputs. q can never leave [0.10, 1.00], so no combination of rows can produce a quality outside what some single mention asserted.
W = Σ wi is carried forward as the evidence weight, and it is the quantity that everything downstream depends on. It is printed on dish pages. Its unit is one fresh, medium-confidence mention: W = 7.5 means “this dish has about seven and a half fresh, unambiguous mentions' worth of evidence behind it”, however many rows that actually took.
Shrinkage: why two raves can't make a ten
This is the most important step on the page, and the one that most distinguishes a Forkelo Score from an average rating.
A weighted mean of two enthusiastic mentions is 1.0 — the same as a weighted mean of two hundred. Publishing that as a 10.0 would be false precision of the worst kind: it would put an under-researched dish above a thoroughly documented one, permanently, on the strength of two comments. So the mean is pulled toward a neutral prior by an amount that depends on how little evidence there is.
Equivalently: ŝ = λ·q + (1 − λ)·m, where λ = W / (W + k) is the share of the answer the evidence has earned. At W = 3 the evidence and the prior are weighted equally.
This is the standard shrinkage estimator — the same device behind Bayesian averages, empirical-Bayes rating systems and the “true Bayesian estimate” used by film and book sites. Read operationally: every dish starts with three units of imaginary, unremarkable evidence in its file, and real evidence has to outvote it.
| Mentions | W | λ | Score |
|---|---|---|---|
| 1 | 1.0 | 0.25 | 6.6 |
| 2 | 2.0 | 0.40 | 7.3 |
| 3 | 3.0 | 0.50 | 7.8 |
| 5 | 5.0 | 0.63 | 8.3 |
| 10 | 10.0 | 0.77 | 9.0 |
| 20 | 20.0 | 0.87 | 9.4 |
Why the prior is 0.55
The prior m = 0.55 is where a dish sits when nothing at all is known about it. It is placed just above the neutral valence of 0.50, and far below the positive valence of 0.85. Three consequences follow, and all three are intended:
- A praised dish is always pulled down. There is no amount of enthusiasm that shrinkage flatters.
- A panned dish is always pulled up. A single bad mention produces a 4.4, not a 1.0 — one disappointed diner does not get to condemn a kitchen.
- A dish with nothing but neutral mentions drifts slightly downward, settling near 5.0 rather than at the prior. Being unremarkable is mildly worse than being unknown.
The half-point above the midpoint reflects a survivorship fact rather than optimism: a dish that a researcher chose to investigate, at a restaurant that is still open, is marginally more likely to be decent than a coin flip. It is a small thumb, and it is the only one on the scale.
Why the prior weight is 3.0
k = 3.0 is the number of units of real evidence needed to weigh as much as the prior. Below W = 3, a dish's own evidence is a minority shareholder in its score. It is set at three because that is roughly the point at which a claim about a dish stops being anecdote — one source is a person, two is a coincidence, three across different places is a pattern.
Note that k is measured in weight, not in rows. Three ancient, low-confidence rows are worth 3 × 0.10 = 0.30 and barely move a score at all; one fresh, high-confidence row is worth 1.20 and moves it four times as much. This is the mechanism by which the system prefers a little good evidence to a lot of bad evidence, and it is why a research agent is told to find the date.
Agreement: the corroboration bonus
Twenty comments in one Reddit thread are, epistemically, close to one source. Four mentions in a newspaper, a food blog, a review site and a Forkelo user's note are four. Independence is what makes evidence add up rather than merely accumulate.
|S+| is the number of distinct source_type values with at least one positive row. The bonus is added after shrinkage and saturates at five source types.
| Independent source types agreeing | Bonus | Points |
|---|---|---|
| 1 | 0.00 | — |
| 2 | 0.02 | +0.2 |
| 3 | 0.04 | +0.4 |
| 4 | 0.06 | +0.6 |
| 5 or more | 0.08 | +0.8 |
Three details in that definition do most of the work.
Only positive rows count toward agreement. A source type that has only ever criticised the dish does not certify it. The bonus measures a chorus, not a crowd.
It is applied after shrinkage, not before. That makes it a genuine addition rather than something the prior can dilute — which matters most for exactly the dishes it is meant to reward.
It is small, and it is capped. Eight-tenths of a point is the entire range. Cross-source agreement is mostly rewarded elsewhere: it is a hard requirement for the confidence label, where two source types are the floor for medium and three for high. The score bonus exists to separate two dishes that are otherwise equal, not to be a second scoring system.
But the cap interacts with shrinkage in a way worth stating plainly, because it is the cheapest route to the top of the scale:
| Target score | With 1 source type | With 5 source types |
|---|---|---|
| 8.0 | W ≥ 3.6 | W ≥ 1.7 |
| 9.0 | W ≥ 9.9 | W ≥ 4.3 |
| 9.5 | W ≥ 21.6 | W ≥ 7.0 |
| 10.0 | W ≥ 267 | W ≥ 12.9 |
A perfect ten is reachable, and it costs about thirteen units of unanimous, superlative, fresh evidence spread across five kinds of source. From a single source type it costs two hundred and sixty-seven, which is to say: it does not happen. That asymmetry is the intended shape of the whole system — breadth of source is worth far more than volume within one.
Scale, clipping and rounding
Clipping. The agreement bonus can in principle push ŝ + a above 1.0; it is clipped before scaling, so 10.0 is a genuine ceiling rather than a number the scale runs past.
Rounding. One decimal place, half away from zero, computed in decimal arithmetic rather than binary floating point. This is not fussiness: banker's rounding would make an 8.45 print as 8.4 on one dish and 8.5 on another depending on the digit before it, and the same evidence must always produce the same printed number. The score is stored at the same precision it is displayed, so what a ranking sorts on is exactly what a reader sees.
The reachable range. Nothing is truly at either end.
| Evidence | W | Score |
|---|---|---|
| One negative mention, nothing else | 1.0 | 4.4 |
| Five negative mentions | 5.0 | 2.7 |
| Twenty negative mentions | 20.0 | 1.6 |
| Nothing but neutral mentions | 20.0 | 5.1 |
| Plain praise only, well evidenced | 20.0 | 8.5 |
| Unanimous superlatives, five source types | 12.9 | 10.0 |
Published scores therefore cluster in a band near the top of the scale, and that is arithmetic rather than generosity: a dish reaches a ranking page only once somebody has researched it, and any dish being praised at all starts from a valence of 0.85 before shrinkage pulls it back. Dishes that would land in the 2s and 3s mostly never get there — a dish with five negative mentions and nothing else is one the researcher is told to stop spending effort on. The bottom of the scale exists so that the top of it means something, not because Forkelo sets out to publish a wall of failures.
The zero that is never published
A dish with no evidence rows returns a score of 0.0 with low confidence. That value is never shown to anybody. When a dish's evidence is emptied, the recompute step deletes the stored score row instead of writing a zero, and the dish keeps whatever score was last asserted for it — or drops out of every ranking if it has none. A 0.0 in this system means “not measured”, and printing it as if it meant “terrible” would be the single most misleading thing the site could do.
Confidence, reported separately
Two facts about a dish are not the same fact: how good it seems, and how sure anyone is. Blending them produces a number that means neither. Forkelo computes them separately and prints them side by side — 9.4 · High confidence reads differently from 9.6 · Early ranking, and it should.
| Label | Requires | Shown as |
|---|---|---|
| high | W ≥ 15 and ≥ 3 distinct source types | “High confidence” |
| medium | W ≥ 5 and ≥ 2 distinct source types | “Medium confidence” |
| low | anything else | “Early ranking” |
The conjunction is the whole point. Thirty Reddit comments and nothing else is still low confidence, because one kind of source is one point of view however many people are speaking from it. Equally, one mention each in three different kinds of source is low confidence too — breadth without depth is as thin as depth without breadth.
Counting in weight rather than rows also means the requirement is honest about quality. Reaching W = 15 takes 13 fresh high-confidence rows, 15 medium ones, or 30 low-confidence ones — and it takes about 24 medium rows if they are all a year old, because year-old evidence weighs 0.630 — or about 38 of them at two years.
Confidence is never folded into the score, and the score is never adjusted to compensate for it. The two numbers travel together everywhere: on the ranking card, on the dish page, in the sort order, and in the API.
A dish at the top of a ranking marked “Early ranking” is a to-do item, not a verdict. It means the evidence is thin, the fix is more research, and the fix is never to adjust the number.
From score to rank
A score is a property of a dish. A rank is a property of a dish within a ranking — a canonical concept in a place, like carne asada burritos in San Jose. Rank is computed in exactly one place in the codebase so that a dish's position is the same number on the ranking page, on its own page, and in the sitemap.
Eligibility
A dish appears in a ranking only if all three hold: it has a score, it is currently available, and its restaurant is not marked closed. A ranking is published as a page only when at least three dishes qualify — below that it is a list, not a comparison, and it is kept out of the sitemap.
Ordering
The sort key is a strict, deterministic tuple:
confidence rank: high = 0, medium = 1, low = 2. Names compare case-insensitively.
Two dishes tied at 9.6 are separated by how well evidenced they are, so a 9.6 marked “Early ranking” never sits above a 9.6 marked “High confidence”. This is the one place where confidence touches ordering, and it is a tie-break, not a weighting — it can never move a 9.5 above a 9.6.
The final two components make the sort total: there is no pair of distinct rows the key cannot separate, so the ranking is stable across page loads, across servers and across re-deploys. Nothing about the reader — location, history, device, whether they are signed in — enters the order. Distance is shown, and never sorted on.
Three dishes, calculated in full
Each of these was produced by running the real scoring function on the rows shown, with the calculation date fixed at 1 September 2026. Every intermediate figure is reproducible.
A. The pair of raves
Two Reddit comments, both today, both superlatives, both read with medium certainty. This is what a dish looks like after ten minutes of research.
| Source | Date | Read | v | ρ | κ | w | vw |
|---|---|---|---|---|---|---|---|
| reddit · positive · strong | today | medium | 1.00 | 1.000 | 1.0 | 1.000 | 1.000 |
| reddit · positive · strong | today | medium | 1.00 | 1.000 | 1.0 | 1.000 | 1.000 |
| Totals | 2.000 | 2.000 | |||||
q = 2.000 / 2.000 = 1.000. Shrinkage: (1.000 × 2 + 0.55 × 3) / (2 + 3) = 3.65 / 5 = 0.730. One source type, so a = 0. Score = 0.730 × 10 = 7.3.
Perfect evidence, and a 7.3 — because there is almost none of it. The prior is doing 60% of the work. This is the system behaving correctly: two people loving something is genuinely weak evidence that it is the best in the city, and no amount of enthusiasm in those two mentions changes that.
B. The properly researched dish
Ten rows across five source types, spanning two and a half years, including one neutral and one negative. This is roughly what a finished research pass looks like.
| Source | Date | Read | v | ρ | κ | w | vw |
|---|---|---|---|---|---|---|---|
| reddit · positive · strong | 2026-07-12 | high | 1.00 | 0.938 | 1.2 | 1.125 | 1.125 |
| reddit · positive · moderate | 2026-03-02 | medium | 0.93 | 0.793 | 1.0 | 0.793 | 0.738 |
| reddit · positive | 2025-11-20 | medium | 0.85 | 0.697 | 1.0 | 0.697 | 0.593 |
| editorial · positive · strong | 2026-05-09 | high | 1.00 | 0.865 | 1.2 | 1.038 | 1.038 |
| blog · positive · moderate | 2026-01-18 | medium | 0.93 | 0.751 | 1.0 | 0.751 | 0.699 |
| review_site · positive | 2026-06-30 | medium | 0.85 | 0.923 | 1.0 | 0.923 | 0.785 |
| review_site · neutral | 2025-08-04 | medium | 0.50 | 0.608 | 1.0 | 0.608 | 0.304 |
| review_site · negative | 2024-09-15 | medium | 0.10 | 0.404 | 1.0 | 0.404 | 0.040 |
| blog · positive · weak | 2024-04-01 | low | 0.88 | 0.327 | 0.5 | 0.164 | 0.144 |
| forkelo_user · positive · strong | 2026-08-21 | medium | 1.00 | 0.986 | 1.0 | 0.986 | 0.986 |
| Totals | 7.490 | 6.451 | |||||
q = 6.451 / 7.490 = 0.861. Shrinkage: (0.861 × 7.490 + 1.65) / (7.490 + 3) = 8.101 / 10.490 = 0.772. Five source types carry positive rows, so a = min(0.08, 0.02 × 4) = 0.080 — the cap. Score = (0.772 + 0.080) × 10 = 8.52 → 8.5.
Worth reading closely. The 2024 negative review contributes only 0.404 of weight, and the 2024 low-confidence blog post only 0.164 — together barely more than half of one fresh high-confidence mention. Meanwhile the agreement bonus is at its cap and is contributing 0.8 points, more than a third of the distance this dish sits above example A. The confidence is medium rather than high despite five source types, purely because W = 7.49 falls short of 15: about eight more fresh mentions would settle it.
C. The single loud source
Thirty fresh Reddit comments, every one a superlative, and nothing anywhere else.
q = 1.000, W = 30.0. Shrinkage: (30 + 1.65) / 33 = 0.959. One source type, so a = 0. Score = 9.6.
This is the case the design is most proud of, and the one that looks strangest. Thirty mentions is plenty of evidence by volume — enough weight for high confidence twice over — and the dish still reads “Early ranking”, because confidence requires three independent source types and this has one. It also forfeits the entire agreement bonus, which is what keeps it below the 10.0 that thirty unanimous raves would otherwise buy.
Compare it with example B: a dish with a quarter of the evidence weight, one negative review and one neutral one, sits only 1.1 points lower and carries a better confidence label. That is the system saying, correctly, that ten mentions in five places tell you more about a burrito than thirty mentions in one.
What actually moves a score
A methodology is only honest if it says how much each lever is worth. These are exact.
The marginal mention
Adding one more fresh, unanimous superlative to a dish that already has n of them (three source types, so a = 0.04):
| Going from | to | Score change |
|---|---|---|
| 1 mention | 2 | 7.0 → 7.7 (+0.7) |
| 2 mentions | 3 | 7.7 → 8.2 (+0.5) |
| 4 mentions | 5 | 8.5 → 8.7 (+0.2) |
| 9 mentions | 10 | 9.3 → 9.4 (+0.1) |
| 19 mentions | 20 | 9.8 → 9.8 (+0.0) |
Sharply diminishing, by construction. The second mention about a dish moves the score about eight times as far as the tenth, and about twenty-five times as far as the twentieth. Research effort is best spent on dishes with two mentions, not on dishes with twenty — which is exactly the behaviour the scoring shape is trying to induce in the people and agents filing evidence.
The cost of one criticism
Adding a single fresh negative row to a dish whose other mentions are all plain praise:
| Existing positive mentions | Before | After | Cost |
|---|---|---|---|
| 3 | 7.0 | 6.1 | −0.9 |
| 5 | 7.4 | 6.7 | −0.7 |
| 10 | 7.8 | 7.3 | −0.5 |
| 20 | 8.1 | 7.8 | −0.3 |
A negative mention is expensive — nearly a full point against a thinly evidenced dish — and gets cheaper as the evidence base grows, which is the same diminishing-returns curve seen from the other side. Note the implication for anyone filing evidence: omitting the one bad review from a set of five is worth about +0.7 to the dish. That is the single largest distortion a careless researcher can introduce, and it is why negative evidence is not optional.
Ranking sensitivity
Because scores are printed to one decimal, the practical question is what it takes to move one place in a ranking. Near the top, where well-researched dishes cluster between 8.5 and 9.5, adjacent ranks are typically 0.1–0.3 apart. In that band, one new source type filing a positive mention (+0.2 from the agreement bonus alone, before its weight contributes anything) is usually worth more than three more mentions in a source type already present.
What the score refuses to use
Just as informative as the inputs. None of the following is read by the scoring function, and several are not stored at all.
| Not an input | Why not |
|---|---|
| The restaurant's star rating | The founding rule. It is an average over everything the restaurant does and contains no information about one dish. |
| Review counts and popularity | A proxy for footfall and marketing budget. A twelve-seat place with a perfect dish would lose to a chain every time. |
| Price | Value is a real question, but it is a different one, and folding it in would make the number mean two things at once. Price is shown, never scored. |
| Distance from the reader | Distance is displayed on request and never sorted on. A ranking must be the same document for everyone who reads it. |
| Anything the restaurant says or pays | Nothing on Forkelo is for sale — not a rank, not a placement, not a tie-break. An owner can correct a fact about their restaurant; they cannot move their dish. |
| Suggestions from visitors | “Tell us what's better” is a research lead, never evidence. It sends a researcher to look; what they find is what gets filed. One anonymous opinion arriving through a form is precisely what Forkelo exists not to rank on. |
| Photographs | Nobody has yet shown that a photograph carries information about a dish's quality that its text mentions do not. |
| Editorial preference | There is no override, no boost field and no manual thumb. A human can add, correct or remove an evidence row, which is auditable; nobody can set a score. |
Decay, drift and recomputation
A score is recomputed from scratch whenever the dish's evidence changes — a row filed, corrected or deleted. There is no incremental update and no cached partial state, so a score can never drift out of step with the rows that justify it.
But there is a second, slower kind of change that happens with no new evidence at all. Because every row's weight decays with age, a dish nobody has mentioned in years loses W, shrinkage pulls harder, and the score slides back toward the prior. A dish does not hold its rank by having once been great.
| Time since the last mention | ρ | W | Score | Confidence |
|---|---|---|---|---|
| now | 1.000 | 10.0 | 9.4 | Medium |
| 1 year | 0.630 | 6.3 | 8.9 | Medium |
| 2 years | 0.397 | 4.0 | 8.5 | Early ranking |
| 3 years | 0.250 | 2.5 | 7.9 | Early ranking |
| 4 years and after | 0.200 | 2.0 | 7.7 | Early ranking |
The dish loses 1.7 points and its confidence label over four years of silence, then stops — the recency floor prevents it from decaying to the prior entirely, because an old rave is not the same as no information. The confidence label falls first and falls hardest, which is the honest ordering: the site becomes unsure about the dish before it becomes negative about it.
Two consequences worth stating. First, a ranking left unresearched does not freeze; it slowly loses its spread as everything in it drifts toward the middle. Second, the ordering of a stale ranking is largely preserved, because every dish in it decays at the same rate — the numbers fall together. What changes is that they all stop claiming to be certain.
Known limitations
Stated plainly, because a methodology page that lists only strengths is marketing.
- The source pool is what is written down. Dishes discussed in English on the open web are better evidenced than dishes discussed in Vietnamese in a private group chat, or not discussed at all. This biases against exactly the kind of restaurant Forkelo most wants to find, and no constant on this page fixes it — only better sourcing does.
- Source-type classification is a judgement call. The agreement bonus and both confidence thresholds depend on it, and the boundary between a blog and an editorial publication is genuinely fuzzy. A researcher who classifies generously inflates confidence; the API cannot detect it.
- Row granularity is not fully standardised. A Reddit thread with fifteen people praising a dish can honestly be filed as one row or fifteen. That is a factor-of-fifteen difference in W, and the largest remaining source of inconsistency in the system. The standing instruction is to be explicit about which was done, and consistent within a dish, but the constraint is procedural rather than enforced by the code.
- The prior is asserted, not estimated. m = 0.55 and k = 3.0 are stated judgements. With enough scored dishes they could be estimated from the distribution of well-evidenced scores — the classic empirical-Bayes move — and at that point they should be. They are not yet.
- Attention is not quality, and the two correlate. A famous restaurant accumulates evidence weight faster than an equally good unknown one, which converts into a higher score through shrinkage even when quality is identical. Every evidence-based ranking has this problem; naming it is not the same as solving it.
- Consistency is not modelled. A dish that is a 10 on a good day and a 4 on a bad one produces the same weighted mean as one that is a steady 7. Variance in the evidence is visible on the dish page as a mixture of praise and criticism, but it does not enter the number.
- The score is not a prediction about you. It measures how strongly, how recently and how widely a dish is recommended. It does not know whether you like cilantro.
Every constant in one table
Nineteen numbers, and there are no others. All of them live in app/services/scoring.py except the ordering and labels, which are noted.
| Constant | Value | Applies to |
|---|---|---|
| positive valence | 0.85 | base value of a positive mention |
| neutral valence | 0.50 | base value of a neutral mention |
| negative valence | 0.10 | base value of a negative mention |
| strength bonus — weak | +0.03 | positive rows only |
| strength bonus — moderate | +0.08 | positive rows only |
| strength bonus — strong | +0.15 | positive rows only, capped at v = 1.0 |
| reader confidence — low | ×0.5 | row weight |
| reader confidence — medium | ×1.0 | row weight |
| reader confidence — high | ×1.2 | row weight |
| recency half-life | 18 months | exponential decay |
| recency floor | 0.20 | reached at ≈41.8 months |
| undated recency | 0.70 | rows with no source date |
| prior quality (m) | 0.55 | shrinkage target |
| prior weight (k) | 3.0 | shrinkage strength |
| agreement bonus per source | 0.02 | per source type beyond the first |
| agreement bonus cap | 0.08 | reached at 5 source types |
| confidence — high | W ≥ 15, types ≥ 3 | both required |
| confidence — medium | W ≥ 5, types ≥ 2 | both required |
| tie-break order | score, confidence, dish, restaurant | ranking_service.py |
Anything you can measure on this site is one of those nineteen numbers applied to a set of evidence rows. There is no twentieth constant, no hidden multiplier, and no model weights.
Version history
| Version | Date | Change |
|---|---|---|
| 1.0 | September 2026 | First published method. Valence, recency-and-certainty weighting, shrinkage toward a 0.55 prior, capped agreement bonus, separately reported confidence. |
Changing any constant on this page changes every score on the site, so changes are versioned here rather than made quietly. When a constant does change, this page is updated in the same commit as the code.
Corrections. If a dish's score looks wrong to you, the useful thing to send is not a different number but a source. Every ranking page carries a “Tell us what's better” form, and it goes to a researcher, not to the score. If a fact about a restaurant is wrong — closed, moved, renamed, the dish is off the menu — write to support@forkelo.com.
Reading further. What Forkelo is and who makes it · Every ranking published so far · What the site stores about you · Terms
Forkelo Score v1 · methodology revised September 2026 · Vorby Studios LLC