Empty Cells · agriculture and food · the tasting practice
The Word for Every Glass
Ilja Croijmans and Asifa Majid published every coded word that wine experts, coffee experts and novices said about five red wines and five coffees, in Dutch and under a CC0 dedication, and that is what makes this check possible. Their printed agreement comes back from those words, 0.0932 against a printed 0.09 for wine experts on wine flavour, and three routes support the aligned reading and reproduce selected printed results, although four of the deposit's twelve response files carry header faults. Then a question the index cannot ask itself: two wine experts' first flavour words match in 9.34 per cent of same-wine person pairs and 8.22 per cent of different-wine response comparisons. Of 24 group comparisons of agreement about the glass, only 2 fall below 0.05, and both put coffee experts ahead on coffee smell.
Two wine experts' first flavour words match in 9.34 per cent of 1,113 same-wine person pairs and 8.22 per cent of 4,453 different-wine response comparisons: 88 per cent of same-wine agreement is already present across wines.
Credit first. Croijmans and Majid deposited their coded transcripts with the Dutch research data archive DANS in 2016, under a CC0 1.0 public domain dedication and together with their own analysis tables. Every count on this page stands on that openness.
Four of the deposit's twelve response files do not mean what their column headers say, and the same layouts are in the depositors' original workbooks, so this is about the spreadsheets, not the archive's conversion. Three independent routes below support the aligned reading and reproduce selected printed results. The printed results stand. A corrected deposit, or a word from the depositors, would settle the four files for certain.
Loading the deposit extract. Every number below is recomputed in your browser once it arrives.
One
Pick a glass
Each glass is one of the five red wines or five coffees the participants described. Within each wine or coffee block, everyone met the drinks in the same fixed order, while the order of the two blocks was counterbalanced; they smelled and then tasted each one, and spoke freely. The authors then coded each description into its main descriptors, in Dutch. Every chip below is one person's first meaningful descriptor for the glass you pick, exactly as coded. Identical words sit together, so agreement is visible before it is counted. The English beside each word is a gloss for the reader, marked where it is the authors' own and otherwise added by this page.
The glass
Waiting for the extract.
The paper measures agreement with Simpson's index, computed the way the deposit's own tables are: for each glass, the sum of n(n − 1) over the words, divided by N(N − 1). That number has a plain meaning. It is the share of same-glass person pairs whose first words are identical. For Wine 3, wine experts, flavour: 22 people make 231 same-glass person pairs, 31 of them match, and the index is 0.1342, which the deposit's table prints as 0.13. The most common word there is tannine, said by 7.
So the index can be asked a question it cannot ask itself: how often do the same people's words match when they describe different glasses? Flip the switch or choose a baseline. For Wine 3 the rate among cross-glass response comparisons is 9.97 per cent. Over all five wines, wine experts' first flavour words match in 9.34 per cent of 1,113 same-wine person pairs and in 8.22 per cent of 4,453 different-wine response comparisons. About 88 per cent of same-wine agreement is already present across wines.
The paper saw part of this, and said so in words:
Interestingly, wine experts described the flavor of all five wines fairly similarly, by using the source-based descriptor fruit ‘fruit’.
Croijmans and Majid, 2016, Results, on the word clouds
The deposit gives the counts behind it. Of 108 first flavour responses from wine experts, the most common were tannine (23), fruit (15) and zuur (14). Of 107 first smell responses: fruit (28), hout (8) and rood fruit (7).
Two
The printed number, from the words
However, when describing the flavor of wine, wine experts had higher agreement (M = 0.09, SD = 0.03) than novices (M = 0.05, SD = 0.02), p = .011, d = 1.56, and coffee experts (M = 0.04, SD = 0.02), p = .007, d = 1.96.
Croijmans and Majid, 2016, Results, Consistency
The page's engine reads the deposited coding and nothing else: rows marked as each person's first response, flavour only, the index for each wine, then the mean and sample standard deviation over the five wines. The printed values are never fed in. It gives 0.0932 (SD 0.0271) for wine experts, 0.0476 (0.0177) for novices and 0.0441 (0.0123) for coffee experts, against printed 0.09 (0.03), 0.05 (0.02) and 0.04 (0.02). All 30 of the 30 wine cells in the deposit's own first-response table agree at two decimals.
The anchor, with its three switches
Waiting for the extract.
Each switch is a real parameter. The with-replacement form gives 0.1352, 0.0949 and 0.0929, which do not match; the deposit's tables hold 3 exact zeros among their 228 cells, and that form can never produce a zero. The article cites Simpson 1949 (reference 61), where the deposit README gives 1954. Counting the last word of each coded descriptor gives 0.0966. Counting the fragments as spoken gives 0.0000: nobody's first words match until they are coded. So the engine is independent of this page's builder but not of the authors' coding, and the switch shows what that coding does.
Reproduction is close, not uniform, and the page says where. Of the 11 first-response means the article prints for wine and coffee, 11 reproduce at the printed precision, and 8 of the standard deviations do. The three that do not: coffee experts on wine flavour (0.0123 against printed 0.02), wine experts on wine smell (0.0584 against 0.05) and wine experts on coffee flavour (0.0483 against 0.02). Counting all responses, each descriptor once per person and glass, 6 of 12 printed wine and coffee means reproduce, and 3 of 3 taste means with their standard deviations. These are observed differences; the page does not guess at their cause.
Three
Four files that do not say what their headers say
The deposit README says of the twelve response files: All of these files are structured in the same way.
Eight are. In four, a column holds something other than its header, in four different ways. A reader who trusts the headers gets strategy letters instead of words for one group's tastes and the wrong first response in three other files. Flip the four files between the two readings and watch what moves.
Two readings of the same cells
Waiting for the extract.
Route one: the files check themselves
TASTES_CoffeeExperts, file 176233
The column headed WordsInDescription holds text, not a count: AnswerCode ends with it in 394 of 394 rows, just as AnswerCode ends with FullResponse in the other two taste files (663 of 663, 566 of 566). The column headed MainResponse holds an S, A, E or O letter in 394 rows, and the column headed SAE holds a counter running down to one in 394. Read in place, FullResponse is the coded descriptor and there is no word count. The deposit's CSV version of this file (176299) has the same layout.
COFFEE_Novices, file 176226
The column headed ResponseOrder holds only zeros and ones, with 188 ones: the first-response flag. The column headed FirstResponse runs down to one in 178 of 178 trials with more than one response. File order is spoken order: all 162 rows marked DoubleAnswer follow an earlier identical answer.
ODORS_Novices, file 176248
The same two columns are exchanged, with the counter running up: the column headed ResponseOrder holds 202 flags, and the column headed FirstResponse counts upward in 174 of 174 trials with more than one response.
COFFEE_CoffeeExperts, file 176238
The flag sits on ResponseOrder one in 198 of 199 trials, 47 of them on a response coded O, for other; 48 flagged rows are coded O in all. In the three wine files and the wine experts' coffee file, 0 are. The README defines the flag as a coding of the first meaningful response (either source or non-source term)
, so the engine takes the first response coded S, A or E.
Each layout is also in the original workbook the depositors uploaded, read offline with openpyxl from the frozen files: 394 letters under MainResponse in the coffee experts' tastes, 188 and 202 flags under ResponseOrder in the two novice files, and 48 flagged rows coded O in the coffee experts' coffee file.
Route two: the deposit's own analysis files
The deposit also holds the authors' strategy coding as separate analysis files. In the odours and tastes file (176300), the coffee experts' taste codes match the in-place reading position by position in 394 of 394 rows. The header reading has no letters to match (0); read as numbers, its counter happens to equal the code in 121. In the wine and coffee file (176308), each trial's marked first response carries the same code as the repaired first response in 195 of 196 coffee-expert trials (151 as headed) and 184 of 188 novice trials (83 as headed).
Route three: the printed statistics
The groups differed in the linguistic strategy used to describe tastes, χ2(4, N = 1496) = 16.91, p = .002, Cramer’s V = .08.
Croijmans and Majid, 2016, Results, Taste naming task
From the strategy file's counts the engine gets χ² = 16.91 with N = 1,496 for tastes and 22.90 with N = 1,698 for odours, as printed. Read by its headers, the coffee experts' taste file holds 0 S, A or E codes, so that test could not have been run on it. Coffee experts' agreement on tastes with all responses is 0.2274 (SD 0.0552) in place, against printed 0.23 (0.06), and 0.4646 by header. Their first-response coffee flavour is 0.0295 (0.0051) with the first meaningful response, against printed 0.03 (0.005), and 0.0126 (0.0080) with the flag as deposited.
Table by table
One rule per table, never tuned per cell. Cells equal at two decimals, declared reading against headers: coffee first responses 22 against 6 of 30, odours 24 against 22 of 30, tastes 21 against 13 of 24; the wine files are aligned and give 30 either way. With all responses the wine table is the loosest, 17 of 30 exact and 30 within one hundredth.
A participant number in one file only
The wine experts' taste file carries 23 participant numbers where their wine file carries 22; the extra one has 22 rows and appears in no other response file. The deposit's taste table reproduces better with that person counted (5 of 8 first-response cells, 7 of 8 all-response) than without (1 and 4). The page keeps the file as deposited.
What would settle the four files for certain is a corrected deposit or a word from the depositors. What this page can say is narrower: read in place, the deposit supports the aligned reading and reproduces selected article results.
Four
Agreement about the glass
The sophisticated objection goes like this. The paper already found that wine experts agree more, that coffee experts do not, and that novices lean on evaluative words; the word clouds showed the rest; and re-counting a study of sixty-three people is noise. The answer is a further result you can operate, with every cell shown and none chosen.
Longer descriptions move the number
The all-response index counts a person's own different words as disagreement, so a group that says more scores lower even if everyone shares the same core word. If G people each give k distinct descriptors and all share exactly one, the index is (G − 1) / (k(Gk − 1)), which falls roughly as one over k squared. For wine flavour, wine experts give 8.71 distinct descriptors per person and glass, coffee experts 6.06 and novices 5.38. The pooled index puts wine experts at 0.0180. The share of pairs with at least one descriptor in common puts them at 78.35 per cent, against 44.42 for coffee experts and 38.86 for novices. The two measures lean opposite ways with length. The paper read the fall like this:
When considering all responses, however, this agreement seems to disappear, possibly because each expert is isolating different components of the wine and coming to a unique linguistic profile for their experience.
Croijmans and Majid, 2016, discussion of the consistency results
Length against the measure, wine flavour, all responses
Waiting for the extract.
Neither level settles it. What holds each person's length fixed is the contrast between the same glass and a different glass: the same people, with the same number of words, are on both sides. For wine experts with all their flavour words, 78.35 per cent of same-wine response comparisons share a word and 78.29 per cent of different-wine response comparisons do.
Every cell, the same glass against a different glass
For each group, drink, task and response set, the engine counts matches between two different people's responses on the same glass and between responses on different glasses. Then, within each person, it shuffles which glass each of their descriptions belongs to, and asks how often the shuffled gap reaches the real one. Each person keeps their own words and their own count. The p is one-sided, and its Monte Carlo error sits beside it.
The shuffle census
The census starts once the extract is read.
At 4,999 shuffles and seed 1: wine experts' first flavour words, 9.34 against 8.22 per cent, p = 0.0760 ± 0.0037 (ten seeds give 0.075 to 0.085), so these data cannot tell that gap from shuffling. With all their words, p = 0.4976 ± 0.0071. The largest gap is coffee experts describing coffee smell with all their words: 27.68 against 19.47 per cent, p = 0.0002 ± 0.0002, the smallest this shuffle count allows. Novices on wine flavour with all their words: 38.86 against 35.05, p = 0.0030 ± 0.0008. Wine experts' first words for coffee smell: 3.03 against 1.41, p = 0.0022 ± 0.0007.
9 of the 24 cells fall below 0.05, several within a Monte Carlo error or two of it. A Bonferroni bound over the twelve cells of one response set is 0.0042; over all of them, 0.0021. At these shuffles 3 cells clear the first bound and 1 clears the second. Wine experts on coffee smell sits on the second bound: at 99,999 shuffles, run by the verifier, it is 0.0014 ± 0.0001 and clears both, while novices on wine flavour is 0.0031 ± 0.0002 and clears only the first.
Between groups
A related question, not tested by the paper's agreement comparisons, is whether one group's words follow the glass more than another's. Permuting group labels among people, the 24 comparisons (eight tasks, three pairs of groups each) give 2 below 0.05, and both are coffee experts on coffee smell with all their words. Their glass-specific gap exceeds the wine experts' by 0.069 (p = 0.0112 ± 0.0011) and the novices' by 0.066 (p = 0.0022 ± 0.0005). Novices against wine experts on wine flavour, all words: a difference of 0.037, p = 0.0707 ± 0.0026, not established. The paper found coffee experts agreeing more on coffee smell only once all responses were counted; this contrast arrives at the same place from another direction.
What counts as a replicate
The article's comparisons treat the five drinks as the replicates, in a mixed ANOVA with Bonferroni corrections that this page does not reproduce. Treat the people as replicates instead, by permuting them between groups (9,999 permutations, uncorrected), and the p-values move in both directions. Wine experts against novices on wine flavour: printed .011; a paired t over the five wines here 0.018; people permuted 0.0487 ± 0.0022, at the 0.05 line. Wine smell, same groups: printed .037, people permuted 0.0052 ± 0.0007. Coffee flavour, coffee experts against novices: printed .237, people permuted 0.0117 ± 0.0011, with the novices agreeing more. This is a different question asked of the same data, not a correction.
The replicate, for the headline comparison
Waiting for the census.
Three cautions travel with every cell. Order: All stimuli were presented in a fixed order within each block
; the wine and coffee block order was counterbalanced, but words that track the glass may still be tracking its place in the sequence. Proxy: a text overlap stands in for the test the authors proposed, by conducting a director-matcher task, where people have to match wines and coffees to descriptions
, and is not that test. Size: five drinks and twenty to twenty-two people per group, so a cell above the line means these data cannot tell, never that nobody tracks the glass.
Five
What the record cannot answer
Press any question. The page answers with the record's own reason.
No question asked yet.
For the wines and coffees there is nothing to score: in the authors' words there is no “correct” answer
. For the odours and tastes, accuracy exists only as the depositors' derived table (file 176283), shown here as theirs: odours 37.7, 38.5 and 36.7 per cent correct for wine experts, coffee experts and novices, tastes 83.8, 82.0 and 76.3, over 22, 20 and 21 participants.
The check
- The files. Each of the six shipped data files used by the engine, hashed in your browser against record.json:
- waiting
- The prose. Figures not yet checked.
- The fast pair counts. Not yet run.
- Every choice, run live. Seven interpretive choices, each a parameter of the engine; the table runs every option.
- Planted faults. Each doctors a copy of the extract (or of the repair manifest) and runs the unmodified engine on it.
No fault planted yet.
- A second implementation. An independent Python program read the raw downloads with its own repair code and its own shuffles. Its same-glass and different-glass rates equal the engine's in 24 of 24 cells, and its p-values sit within 1.84 combined Monte Carlo errors of these.
- Free choices, named. Which reading of the four files (declared repairs, each shown above); the comparison rule (two decimals, rounded half up, or the printed precision for printed means); the all-response rule (each descriptor once per person and glass); the index form; the descriptor; the baseline; the measure; the replicate; the shuffle count and seed. English glosses drive no count.
- Not modelled. The ANOVA models; order effects; perception, memory or skill; any accuracy the depositors did not derive; any language but Dutch; a director-matcher test.
- Personal data. The deposit's personal-data status is recorded as unknown, so this page ships a minimised extract: no transcripts, no timings, nothing from the participant file, and participant numbers re-keyed within each group with a stable pseudonymous key across response files for the within-person shuffle. One town name is replaced in a spoken fragment (2 cells). The extract keeps 11,982 response rows; 6,690 wholly blank rows served by the file endpoint are dropped.
- Novelty, bounded. We searched PLOS ONE, PubMed Central, Crossref, web searches for corrections, reanalyses and replications of Not All Flavor Expertise Is Equal, Croijmans and Majid's later work on the language of wine experts, director-matcher and communication-accuracy tests of these descriptions, and the Artificial Wasteland corpus on 2026-09-14 and did not find a same-glass against different-glass comparison of these descriptions, a participant-level permutation of the printed agreement contrasts, or a notice of the header faults in four of the deposit's twelve response files. Crossref records no correction or update for the article.
- Sources and licences.
- Croijmans, I. and Majid, A. (2016). Not All Flavor Expertise Is Equal: The Language of Wine and Coffee Experts. PLoS ONE 11(6): e0155845. doi:10.1371/journal.pone.0155845. CC BY 4.0; short passages quoted with attribution.
- Majid, A. and Croijmans, I. (2016). Human olfaction at the intersection of language, culture and biology. DANS. doi:10.17026/dans-zke-2wgq, version 2.1, 45 files, retrieved 2026-09-14. CC0 1.0; changes: blank rows dropped, transcripts and timings omitted, participants re-keyed within each group with the stable key retained across response files for group-level calculations, one place name replaced. The wine and coffee names are from its Methods section, Table 2 (file 176237).
- Prior work on the same data and questions: Croijmans and Majid (2015), odour naming in the CogSci proceedings, pages 483 to 488; Majid and Burenhult (2014), Cognition 130, the lineage of the index; Brown and Lenneberg (1954), codability; Lantz and Stefflre (1964), communication accuracy, the classic answer that a glass contrast approximates (not re-read for this page); Croijmans, Hendrickx, Lefever, Majid and van den Bosch (2020), Natural Language Engineering 26(5), 511 to 530, a different corpus of wine reviews; Croijmans's doctoral dissertation, Radboud University (2018, not read); Casillas, Rafiee and Majid (2019), odour naming by Iranian herbalists and cooks.