Skip to content

Eleven stages, explained twice

How the model works

Eleven stages, from a PDF file to a matrix a nurse can use. Each stage is explained twice: once in everyday language, once in its technical terms.

  1. 2.1b

    Domain relevance filter

    Is this article actually about palliative care and drug side effects? If not, it stops here.

    Technical

    A relevance score over six term groups; threshold 8. Of 405 articles, 263 passed.

  2. 2.2

    PDF text extraction

    The PDF becomes plain text, page by page.

    Technical

    PyMuPDF, `page.get_text('text')` joined across pages.

  3. 2.3

    Preprocessing

    Tidying the text: words broken at a line end are rejoined, citation numbers are lifted out of the sentence.

    Technical

    NFKC, Unicode dash folding (U+2013 appears in 8.6% of sentences), superscript citation-marker separation.

  4. 2.4

    Sentence segmentation

    The text is cut into sentences. Very short ones are usually layout debris; very long ones are usually a table that turned into prose.

    Technical

    NLTK punkt, then word-break repair against a corpus vocabulary, then a 30–2,000 character length gate. 137,906 sentences, 92,713 into NER.

  5. 3.1

    Entity tagging

    Two AI models read each sentence and highlight which words are drug names and which are side effects.

    Technical

    BioBERT chemical + diseases side by side, batch 8, 512 tokens, plus dictionary matching over 38/32 terms.

  6. 3.2

    Term normalisation

    One thing can be written many ways. All of them collapse to a single standard term.

    Technical

    Lemmatisation, British→American spelling, RapidFuzz with two guards: refusing a mapping to a more specific term, and refusing drug pairs that merely look alike.

  7. 3.2b

    Strict entity filter

    Fragments that got through as entities but are not medical terms are dropped here — one- and two-letter abbreviations, numbers, and ordinary words.

    Technical

    A letter-count threshold, a list of 32 disallowed terms, and an exemption for the suffixes typical of drug names. Its counters report in, out, and dropped.

  8. 3.3

    Co-occurrence (baseline path)

    Every drug is paired with every side effect appearing in the same sentence. No meaning has been checked yet.

    Technical

    Token distance ≤30. Negation is detected over the side-effect span (NegEx via negspaCy), not the whole sentence; negated pairs are weighted 0.65 rather than discarded.

  9. 3.4

    Association scoring

    The more often a pair appears together, and the rarer each is alone, the stronger the relation is taken to be.

    Technical

    PMI over weighted frequency; Final Score = PMI · log(1+frequency); Confidence = sigmoid.

  10. 4.1

    Advanced relations and meaning filters

    This is where the difference lies. The sentence structure is examined: is the drug really the cause, does the sentence deny it, is the drug in fact being used to treat that effect, and does the effect actually belong to a different drug in the same sentence.

    Technical

    spaCy dependency parsing with a causal co-occurrence fallback. Recorded rejections: NO_CAUSAL_SIGNAL, NON_ADVERSE_CONTEXT, DISTANCE_EXCEEDED, attribution to another drug.

  11. 4.1b

    Frequency gate

    A pair only enters the operational matrix if it appears at least twice across the whole corpus.

    Technical

    `advanced_min_pair_frequency = 2`. Lowering it to 1 raises recall 0.800→0.933 but drops precision 0.632→0.538.

What this site cannot do

Three limitations that have to be stated

1. The frequency gate is corpus-level

The operational matrix requires a pair to appear at least twice across all 405 articles. A document you upload almost never satisfies that on its own, so the “passed the gate” figure in the simulation is not a judgement on whether that pair belongs in the matrix.

2. The word-repair vocabulary comes from the corpus

Rejoining broken words (“consti pation” → “constipation”) uses a word list built from the 137,906 sentences of the research corpus. A document from outside the corpus therefore behaves slightly differently than it would had it been part of it. This is a structural limit, not something that can be measured away.

3. Two numeric regimes, measured at zero difference

The article's numbers were computed on GPU fp16; this site runs on CPU fp32. The two were tested on three separate axes — extraction fidelity (GPU), precision (CPU), and batch composition — and all three came back at zero difference across 211 advanced pairs and 333 baseline pairs. So what you see here does not approximate the article's numbers; it is the same numbers.