# Automated leech detection

**URL:** <https://forums.ankiweb.net/t/automated-leech-detection/56887>\
**Category:** FSRS\
**Created:** [March 8, 2025, 3:28pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887 "2025-03-08T15:28:36Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [March 8, 2025, 3:28pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/1 "2025-03-08T15:28:36Z")

</div>

Given the card’s history, we can either store or re-calculate the probability of recall predicted by FSRS, and then use the [Poisson binomial distribution](https://en.wikipedia.org/wiki/Poisson_binomial_distribution) to calculate the probability of a given number of successes.

> **[Discord - Group Chat That’s All Fun & Games](https://discord.com/channels/368267295601983490/1282005522513530952/1347949095012270152)**
>
> Discord is great for playing games and chilling with friends, or even building a worldwide community. Customize your own space to talk, play, and hang out.

> **[GitHub - tsakim/poibin: Poisson Binomial Probability Distribution for...](https://github.com/tsakim/poibin)**
>
> Poisson Binomial Probability Distribution for Python

> I am not even going to try to understand the math with complex numbers, but the usage is actually fairly simple. You just give it a list of probabilities for each trial and the number of successes, and then you can calculate the probability of this many successes _or even fewer_.  
> Example: `p = np.asarray([0.9, 0.85, 0.95, 0.92, 0.87])`  
> `n_succ = 2`  
> This gives me a p-value of 0.836%. So if a card has been reviewed 5 times with these probabilities (note that the order doesn’t matter) there is a 0.836% chance that 2 _or fewer_ reviews will be successful.

We can identify leeches with as few as two or three reviews!  
For example, if the probabilities of recall are 90%, 92% and 93%, then the probability of getting the card right zero times is 0.056%. The probability of getting the card correct once _or_ zero times is 1.948%.  
The higher desired retention, the higher the probabilities, the faster we can identify leeches. At DR=95% we can identify leeches with merely 2 reviews! Btw, the probability of 0 successes if both reviews had a 95% of success is 0.5%.

Btw, I think 1% is a reasonable cutoff. If a card has been failed so much that the chances of it happening (or having even more fails) normally are \<1%, I think it’s most likely a leech.

EDIT: I came up with a good way to correct the threshold to ensure that we don’t tag too many cards, but that is beyond the scope of this topic.

There are 2 challenges:

1. Implementing this mathematical function in Rust.
2. Storing or re-calculating R for every review.

Then we can add a “Automatic leech detection” button here as an alternative to “Leech threshold” when FSRS is enabled.

 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/2/0/2076287128b0bcbbe686d8ab78e9e6542acfa9fe.png)

Now the big question is: do we want a “Recalculate leeches” button if automatic leech detection is enabled? 🤔  
Since changing FSRS parameters will change retrievability at the time of the review, which in turn can change whether the card counts as a leech or no.

@L.M.Sherlock

Also, I asked Claude 3.7 Thinking to re-write it in Rust and remove the calculation of p-values (I calculate them from the PMF) and CDF, leaving only PMF. Idk if it’s any good, but so far Claude 3.7 Thinking has been really freaking good, at least for Python.

> **[PoiBin.rar](https://drive.google.com/file/d/10QaOXwyh8F58wRTlGizOUc0VOaEOqBIy/view?usp=sharing)**
>
> Google Drive file.

Btw, [this repo](https://github.com/tsakim/poibin) uses Fast Fourier Transform to calculate the probabilities approxiamtely, but me and Alex tested using the exact (combinatorics) method and found that for n=64 reviews it’s fast enough that we don’t need FFT, so the Rust code uses the exact approach.

---

<div class="post-metadata">

**Author:** ![sound](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/sound/32/25233_2.png) [@sound](https://forums.ankiweb.net/u/sound)\
**Post date:** [March 8, 2025, 3:30pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/2 "2025-03-08T15:30:57Z")

</div>

That’s really a very good idea way more fitting than the old one (counting lapses) for FSRS

---

<div class="post-metadata">

**Author:** ![L.M.Sherlock](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/l.m.sherlock/32/16770_2.png) [@L.M.Sherlock](https://forums.ankiweb.net/u/L.M.Sherlock)\
**Post date:** [March 12, 2025, 5:43pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/3 "2025-03-12T17:43:03Z")

</div>

Does the detector consider the same-day reviews?

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [March 12, 2025, 5:45pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/4 "2025-03-12T17:45:06Z")

</div>

No, since we need to use probabilities predicted by FSRS.

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [March 13, 2025, 11:30am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/5 "2025-03-13T11:30:21Z")

</div>

@dae I want to bring your attention to this  
Right now we have two ideas for a new leech detector: this (with a little bit of extra math and rules not mentioned in this topic) and a machine learning based detector. The latter would require thousands, if not hundreds of thousands of **manually** labeled (leech/not a leech) cards, so that is not going to happen.

The problem with my idea is that we won’t know how well it works until we try it. Jarrett cannot implement it in the Helper add-on first.  
 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/2/a/2a6b4748205bf890c67cae6d474174fcce093747.png)

And he doesn’t want to do all the work of implementing a leech detector only for you to not merge the PR.  
 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/a/c/ac04120e23ef6f5912a454a69541c2d9ea866cd5.png)

So ideally I’d like you to say “Sure, we can test this idea in a beta and/or as an experimental feature that can be removed if it doesn’t work well”, and then Jarrett would (hopefully) feel motivated to do it.

EDIT: @rossgb implemented it (not in Anki itself): [GitHub - rbrownwsws/leechkit](https://github.com/rbrownwsws/leechkit)  
We’re currently testing it

---

<div class="post-metadata">

**Author:** ![A\_Blokee](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/a_blokee/32/17955_2.png) [@A\_Blokee](https://forums.ankiweb.net/u/A_Blokee)\
**Post date:** [March 13, 2025, 11:03pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/6 "2025-03-13T23:03:53Z")

</div>

I had an attempt at graphing the cards’ percentages:

> <https://github.com/Luc-Mcgrady/Anki-Search-Stats-Extended/pull/36>
>
> Plots the probability that that cards have the given amount of lapses.
> 
> https:…//forums.ankiweb.net/t/automated-leech-detection/56887
> Thought it might work as a graph :shrug:.
> 
> !\[image\](https://github.com/user-attachments/assets/d39d065b-2854-4392-8ea2-32c1fb71ef6a)
> !\[image\](https://github.com/user-attachments/assets/a04c3876-c5f7-4f6e-9b4c-2fe85555c712)
> 
> My code might just be bad but it doesn't seem to work very well at finding leeches.
> Appears after you run the memorised graph.
> 
> I don't really want to take any ownership of this issue so I will probably just merge this as a bad graph.

Here’s how I implemented it if someone wants to check it:

> <https://github.com/Luc-Mcgrady/Anki-Search-Stats-Extended/blob/70d864e95d42faecda821fc53da3559dd6c0532d/src/ts/MemorisedBar.ts#L153-L162>

Doesn’t seem to work well in its current form.

 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/f/4/f46abdcaae35f8af9fa1f95ad5df8417480302c6.png)

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [March 13, 2025, 11:16pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/7 "2025-03-13T23:16:09Z")

</div>

Are you sure you aren’t using same-day reviews and the first review?

Also, I can’t verify the code, so idk if it’s implemented correctly. I guess you can give me probabilities for a given card and the output of your function, and I’ll see if it matches mine

---

<div class="post-metadata">

**Author:** ![A\_Blokee](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/a_blokee/32/17955_2.png) [@A\_Blokee](https://forums.ankiweb.net/u/A_Blokee)\
**Post date:** [March 13, 2025, 11:18pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/8 "2025-03-13T23:18:30Z")

</div>

> <https://github.com/Luc-Mcgrady/Anki-Search-Stats-Extended/blob/70d864e95d42faecda821fc53da3559dd6c0532d/src/ts/MemorisedBar.ts#L153-L154>

> <https://github.com/Luc-Mcgrady/Anki-Search-Stats-Extended/blob/70d864e95d42faecda821fc53da3559dd6c0532d/test/memorised.test.ts#L104-L117>

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [March 13, 2025, 11:19pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/9 "2025-03-13T23:19:19Z")

</div>

Can you give me a list of probabilities for any card, and what your function outputs for those?

I don’t know TypeScript or any Anki-specific code

---

<div class="post-metadata">

**Author:** ![A\_Blokee](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/a_blokee/32/17955_2.png) [@A\_Blokee](https://forums.ankiweb.net/u/A_Blokee)\
**Post date:** [March 13, 2025, 11:22pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/10 "2025-03-13T23:22:32Z")

</div>

Sorry this is the best I can do:  
I think I forgot to multiply the percentages by 100 😅

50% for this review history:

 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/b/3/b383cc99d9f00ccd145dc6fde07429f8fba71470.png)

 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/e/7/e70cd9bea2043fb2e841d2adafa0f6c8ea76941f.png)  
 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/2/5/25a60a9f91069599c3cbfd11e23234158ae9c3bd.png)  
 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/f/c/fce6c83b12b969daf8cb33d4516c4bd5b8ee0ae5.png)

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [March 13, 2025, 11:25pm UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/11 "2025-03-13T23:25:53Z")

</div>

![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/5/0/50561261aa048d0a0f33922d257ef3557a0d0ca5.png)  
Seems about right

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [March 14, 2025, 12:30am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/12 "2025-03-14T00:30:10Z")

</div>

Btw, Jarrett said “Anki doesn’t provide the API to calculate the historical retrievability”. Maybe you can help him? Then we could try the leech detector out using the Helper add-on.

---

<div class="post-metadata">

**Author:** ![A\_Blokee](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/a_blokee/32/17955_2.png) [@A\_Blokee](https://forums.ankiweb.net/u/A_Blokee)\
**Post date:** [March 14, 2025, 1:25am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/13 "2025-03-14T01:25:27Z")

</div>

> <https://github.com/open-spaced-repetition/fsrs-rs/pull/302>
>
> This PR adds a new method \`historical\_memory\_states\` to the FSRS implementation,… which allows retrieving the memory states after each review in a card's history. It also refactors the tensor creation logic into a separate helper function \`item\_to\_tensors\` to reduce code duplication.
> 
> \## Changes
> 
> 1. Extracted tensor creation logic into a reusable helper function \`item\_to\_tensors\`
> 2. Added new \`historical\_memory\_states\` method to retrieve memory states after each review
> 3. Updated the \`infer\` method to use the new helper function
> 
> \## Benefits
> 
> It's the prerequisite to reduce the Algorithm Complexity from O(n^2) to O(n) in card stats of Anki:
> 
> https://github.com/ankitects/anki/blob/9b5da546be49f37c8d6c286e09c86074b2f0c278/rslib/src/stats/card.rs#L145-L160

I think he’s got that covered.

---

<div class="post-metadata">

**Author:** ![ran9](https://avatars.discourse-cdn.com/v4/letter/r/e19adc/32.png) [@ran9](https://forums.ankiweb.net/u/ran9)\
**Post date:** [March 15, 2025, 6:38am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/14 "2025-03-15T06:38:55Z")

</div>

> [@Expertium](#):
>
> We can identify leeches with as few as two or three reviews!

User would avoid wasting time on failing. When user is alerted they can improve the card or put more effort in or drop it. This kind of like the thing the first paragraph should say for us layman.

---

<div class="post-metadata">

**Author:** ![Evelf](https://avatars.discourse-cdn.com/v4/letter/e/ea5d25/32.png) [@Evelf](https://forums.ankiweb.net/u/Evelf)\
**Post date:** [April 8, 2025, 1:09am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/15 "2025-04-08T01:09:40Z")

</div>

> The higher desired retention, the higher the probabilities, the faster we can identify leeches.

I don’t get this part..

My reasoning is that if the DR is higher, the card will be shown sooner, so the probability of getting is wrong is lower. What makes you think otherwise?

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [April 8, 2025, 10:47am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/16 "2025-04-08T10:47:30Z")

</div>

If a card is a leech, it will be failed more often than FSRS predicts. That’s how we define leeches with the new detector. So yes, the probability of recall will be higher at higher DR, but since leeches have a lower p(recall) than FSRS predicts, they will be failed more often. So depending on how much lower it is exactly, it’s possible that leeches can be identified faster at higher DR because you will do reviews more frequently, so the necessary information for the detector will be gathered faster.

Anyway, @dae, we figured out all the math, and I already wrote a detailed specification of the leech detector for @jakep (which he may or may not decide to implement, lol), but there is a problem. The current leech **tag** works on a per-note basis, meaning that all siblings get tagged as leeches. This is undesirable, and even more undesirable with the new detector. So we need to use a **flag** instead of a tag for marking individual cards as leeches. Ideally, both the old lapse count detector and the new one should use a flag instead of a tag.

(writing this made me realize how confusing the whole tag vs flag thing is)

@L.M.Sherlock

---

<div class="post-metadata">

**Author:** ![vaibhav](https://avatars.discourse-cdn.com/v4/letter/v/c68b51/32.png) [@vaibhav](https://forums.ankiweb.net/u/vaibhav)\
**Post date:** [April 8, 2025, 11:13am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/17 "2025-04-08T11:13:10Z")

</div>

> [@Expertium](#):
>
> So we need to use a **flag**

Do we really need to use flags or tags? Can’t we just store the p value in the data column (just like FSRS memory states) and then Anki would show all cards having p \< 0.05 (or whatever you choose) when searching `is:leech`?

Just like FSRS memory states, the stored p value will be updated on each review and each time the FSRS parameters are updated.

This approach will also allow the user to search cards by defining their own threshold for the p value.

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [April 8, 2025, 11:17am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/18 "2025-04-08T11:17:39Z")

</div>

Then we have to explain probabilities and whatnot to users instead of a simple “This card is a leech” thingy. I’m trying to keep things simple. Btw, that also includes **no** settings for the automatic leech detector. It will be a black box with a toggle. We’ll need to choose a leech threshold, like 1% or 2% or 5%.

 ![image](https://us1.discourse-cdn.com/flex002/uploads/anki2/original/3X/6/6/66f9296515402df12fbbdb9b0daf04e9a6c5fe7a.png)

We can make it possible to search for cards based on their p(leech) so that power users can do power user things, but the detector should work purely automatically, so that most users can just turn it on and forget about it.

---

<div class="post-metadata">

**Author:** ![vaibhav](https://avatars.discourse-cdn.com/v4/letter/v/c68b51/32.png) [@vaibhav](https://forums.ankiweb.net/u/vaibhav)\
**Post date:** [April 8, 2025, 11:20am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/19 "2025-04-08T11:20:39Z")

</div>

You don’t have to explain anything to a basic user.

- The basic user would search `is:leech` in the Browser and Anki will “internally” search for cards with p \< 0.05
- The advanced user can search for `prop:leech-p<0.01`

---

<div class="post-metadata">

**Author:** ![Expertium](https://sea2.discourse-cdn.com/flex002/user_avatar/forums.ankiweb.net/expertium/32/24176_2.png) [@Expertium](https://forums.ankiweb.net/u/Expertium)\
**Post date:** [April 8, 2025, 11:21am UTC](https://forums.ankiweb.net/t/automated-leech-detection/56887/20 "2025-04-08T11:21:51Z")

</div>

That’s pretty much what I’m saying. The detector adds and removes the leech flag automatically, and the user doesn’t need to think about it. If the user wants to think about it, he can do something like `prop:leech-p<0.01`

If we don’t use flags/tags to mark leeches, that’s a net loss of functionality

[Next page](https://forums.ankiweb.net/t/automated-leech-detection/56887.md?page=2)
