I have no problem with that. But then I clarify but you still persist in your assumptions about me. What I asked Expertium simply meant “if you already know the result of this experiment why you tacitly approving the idea of 2 button use being better AND how this changes anything”. It’s always more humble to assume you’re in the wrong especially when talking about the intent of another person. I know myself better than you.
Right here. But I don’t think we should further distract this text if you want to continue discussing, we can do it elsewhere.
I answered why I pinged you, I was happy to clarify. But I think you didn’t understand what I was saying maybe?
I’m not sure what assumption you think I made about you, I didn’t make any such assumption that you’re thinking. What I said, in other words, is that you’re interested in knowing results of 2-button use vs. 4-button use.
But let me answer this question, which is complicated so I hope my answer is clear - “if you already know the result of this experiment why you tacitly approving the idea of 2 button use being better AND how this changes anything”
Knowing the result in advance is actually not important compared to knowing the actual numbers from the result. It very well changes things dramatically, bc it could demonstrate equivalency of 2-button use vs. 4-button use, not 2-button users vs. 4-button users.
So, from 21K users, only 6.7K have at least a card with an interval>730?
Why the threshold ends at 0.500 (half-life?)
The higher the threshold the closer the RMSE, does that mean that for long intervals 2 buttons vs 4 buttons becomes less relevant?
That’s not how the English language works. You say a non-sensical sentence and I am supposed to not question it? Why say something if others are not allowed to ask for clarification? Yes I agree I was a bit rude in that. But that is because you made a very definitive statement about what I wanted.
Maybe I don’t know enough English to understand then I apologise.
Is it intentional that you keep saying things like “you’re interested…” “you want to…” No sir I completely understand everything you think I don’t understand.
I think you don’t get my point. For that particular thing I wrote I was interested in what Expertium was thinking. That question also didn’t mean what you previously thought it meant.
from 21K users, only 6.7K have at least a card with an interval>730?
From 20k users, but yes.
Why the threshold ends at 0.500 (half-life?)
The threshold has nothing to do with half-life, I don’t even know what you mean.
The higher the threshold the closer the RMSE, does that mean that for long intervals 2 buttons vs 4 buttons becomes less relevant?
No, above 50%, the curves become meaningless since you basically say, “If someone barely uses Hard and Easy, then he’s a four button user, and if someone uses Hard and Easy a lot, he’s a two button user," which makes no sense. There is just no reason to use thresholds >50%. And intervals have nothing to do with this, aside from just filtering out users with max(delta_t)<730. You seem very confused.
The X axis, which ends at 0.5,
Then what is delta? I thought 730 were days
delta means difference in simple words. delta t generally means time interval. probabaly that.
The threshold one I don’t understand. because Expertium said
That means if threshold is .6 (>50%) then Hard+Easy button use have to be > than 60%. I don’t understant what “If someone barely uses Hard and Easy then he’s a four button user” mean.
I think we should clarify what “better” means here. 4-button users have high RMSE on average means FSRS is not good at predicting the stability when users are pressing “hard” or “easy”. But it doesn’t mean 4-button is worse than 2-button.
Some question may clarify this issue:
- Does 2-button users spend less time to remember more cards than 4-button users?
- Is true retention of 2-button users closer to their desired retention than 4-button user?
Maybe the more you press Hard and Easy FSRS performs less well at predicting the stability for every button. Isn’t that possible?
I thought RMSE is telling us just that.
That’s a really good question to investigate. I assume we try to find workload:knowledge for that. The lower the better.
It’s still a problem of the model. @Expertium could you run the analysis on the result of GRU?
It doesn’t because the dataset is collected from users who use SM-2. SM-2 doesn’t have desired retention.
Yes I was aware of that. But with SM2 I imagine a lot of these inquiries would be meaningless as how well SM2 performs depends on how you configure the settings.
I actually suggested (along with default 2 buttons) that FSRS be made the default scheduler of Anki.
My bad, I was dumb. Disregard what I said.
X axis is threshold. It’s a proportion. Delta_t is interval length, but it’s not on the graph. I just removed users with maximum interval length <=730 days.
Not really. RMSE tells you how well the algorithm predicts probability of recall on historical data. True Retention tells you how well it works in practice when it’s actually used for scheduling.
Will do. And I’ll do DASH too, I’m curious since IIRC DASH doesn’t differentiate between Hard/Good/Easy, but still performs remarkably well.
It would be a bit weird if you call people who use Easy+Hard 95% of the times 4 button users. Let’s call them Easy/Hard users. Anyways, are there collections in the dataset where threshold is above .5? I would be interested to know how graph looks like in that region.
@L.M.Sherlock @sorata @suiyuan @nmjkjm I did the same analysis for GRU (a neural network) and DASH (another model of human memory, but it doesn’t differentiate between Hard/Good/Easy).
Green area indicates statistical significance.
You can see that the pattern is the same. Conclusion: four buttons suck.
For some people this is not making intuitive sense but evidence is evidence. Maybe for FSRS-17 it won’t make a difference (/s) but it makes a difference now. Also I don’t think there are other relevant differences between these two groups that will be affecting RMSE.
But @Expertium do you think there will be difference in average workload:knowledge/knowledge acquisition rate for 4-button users/2-button users?
People spend more time on "Hard than on “Good”, and less time on “Easy” than on “Good”, so idk.

