A tool earns its place when the recurring benefit is larger than the recurring cost, and the recurring cost is bigger than you think, because it includes maintaining the configuration, the tax of a second place to look, and any new kind of mistake the tool makes possible. You cannot settle this by reflecting on it. Fourteen days of rough measurement will, and the last four days matter most, because that is when you find out whether you still use it once you have stopped trying.
This is about a tool you have already adopted. If you are earlier than that and trying to work out which of your recurring tasks deserves attention at all, the seven day friction audit is the thing to run first.
- Some of a new tool's apparent benefit is just the attention it is getting.
- The more setup time you invested, the more kindly you will rate it. Measure instead.
- Recurring costs decide the answer, and recurring costs are the ones people omit.
- The durability phase is the whole point. A tool you use only when reminded saves nothing.
- Modify is the most common right answer and the least chosen.
Why you cannot just think about it
Ask somebody whether a tool they adopted last month is helping and they will answer immediately and confidently. There are two specific reasons that answer is unreliable.
The new thing gets attention the old thing lost. When you adopt a tool you also start paying deliberate attention to a process you had been running on autopilot for years. Some of the improvement is real and some of it is the attention. These feel identical from the inside, and only the first survives the attention going away.
Effort spent makes you generous. The afternoon you lost to configuring it does not disappear. It quietly raises your estimate of how useful the result is, because the alternative is concluding the afternoon was wasted. This runs in the same direction for everybody and is the reason to write numbers down rather than consult your impression.
The question is not whether the tool is good. It is whether the version of you who is busy and distracted still comes out ahead.
Neither of these is a character flaw and neither goes away with awareness. They are just reasons to measure something cheap instead of trusting a recollection.
The costs that get left out of the sum
Everybody counts setup time, because it hurt. Almost nobody counts the recurring costs, and those are the ones that decide the answer over a year.
| Cost | What it looks like | Usually counted? |
|---|---|---|
| Setup | Initial configuration, import, learning | Yes |
| Maintenance | Keeping settings, templates and rules current | Rarely |
| Switching | Moving attention in and out of another window | Almost never |
| Second place to look | Checking two systems because work now lives in both | Almost never |
| New errors | Mistakes that were not possible before | Rarely |
| Migration drag | Old work still in the old place, indefinitely | Rarely |
| Recovery | Getting back on track after not using it for a week | Almost never |
The two worth dwelling on are the ones nobody counts.
A second place to look is expensive out of proportion to its size. If half the work lives in the new tool and half is still in the old one, every question now requires two lookups and a judgement about which is authoritative. That cost is paid on every single retrieval, forever, and it is invisible because each instance is a few seconds.
New errors are worse than slow work. A tool that saves ten minutes a day and creates one wrong figure a month may be a bad trade, depending entirely on what happens downstream of the wrong figure. Time saved and errors created are not in the same units and cannot be netted off casually.
The sum, kept honest
Write it as arithmetic on your own measurements rather than a score out of ten.
Minutes saved per week, minus maintenance minutes per week, minus switching minutes per week, minus the weekly share of setup spread over the time you expect to keep using it. Then, separately and not converted into minutes, the change in errors.
Two deliberate choices in that formulation.
Setup is spread across the period you expect to keep the tool, which stops a single bad afternoon from dominating a decision about the next year, while still counting. If you cannot say how long you expect to keep it, that is itself informative.
Errors stay out of the minutes total. Converting a mistake into a time cost requires an exchange rate you would be inventing. Report it alongside instead: "saves about forty minutes a week and has caused two wrong figures" is a more honest sentence than any single number, and it is easier to decide with.
The fourteen day test
Three phases, and the third is the one people skip and the reason the whole thing works.
Days one to three: baseline. Do the task the old way, or note how long it took before if the tool is already embedded. You need a number to compare against, and estimating it afterwards is exactly the recollection problem you are trying to avoid. Record time per instance and how many instances happened.
Days four to ten: deliberate use. Use the tool properly and consistently. Fix the configuration if it is wrong. This phase measures the tool at its best, which is the fair comparison. Record the same two numbers, plus maintenance time and anything that went wrong.
Days eleven to fourteen: durability. Stop making an effort. Do not remind yourself to use it, do not tidy its settings, do not open it out of duty. Then record what actually happened.
The durability phase answers the only question that matters over a year: does this survive contact with a normal week. A tool that helps when you are focused on it and gets abandoned when you are busy has a real value close to zero, because the weeks when you are busy are the weeks you needed help.
What to record
Six fields, one line per instance. Same principle as any measurement you actually want to complete: if a line takes longer than fifteen seconds, you will stop.
- Date and phase. Baseline, active or durability.
- Minutes. Rounded to five. You are recalling, not timing.
- Used the tool? Yes or no. In the durability phase this is the important column.
- Prompted or spontaneous? Did you reach for it, or did something remind you.
- Maintenance minutes. Time spent on the tool rather than on the work.
- Anything wrong. Errors, rework, confusion, a result you had to check twice.
Two weeks of this is perhaps ten minutes of writing in total. It is not a study and does not need to be. It needs to be better than your memory, which is a low bar and the entire point.
Continue, modify, or remove
Continue when net weekly value is clearly positive, errors have not increased, and you kept using it during the durability phase without prompting. Then stop thinking about it. A tool that passes does not need re-examining every quarter.
Modify when the tool does something useful but the numbers are marginal. This is the most common correct answer and the least often chosen, because it requires more thought than either keeping or deleting. Typical modifications:
- Use it for one kind of work rather than all work, ending the two places to look problem.
- Move it earlier or later in the sequence so it stops duplicating a step.
- Turn off the features that generate maintenance, which is usually most of them.
- Stop migrating old work into it and let the old place age out.
Remove when net weekly value is negative, or when the durability phase shows you did not reach for it once. Removal is cheap and reversible: the tool still exists and you can adopt it again when the situation changes. Keeping a tool that does not pay compounds, because it accumulates configuration, data and habit until leaving becomes a project.
Removing a tool costs an afternoon. Keeping the wrong one costs a few minutes every day, quietly, for years.
A worked example
Take a common case: someone adopts a note taking application to replace a folder of text files.
Baseline shows finding a note took under a minute, most days, several times a day. Unremarkable, which is why nobody had ever questioned it.
The active week looks good. Search is faster, the mobile app means notes get captured that previously did not, and the whole thing feels like an obvious improvement. If the test ended here the answer would be continue.
The durability phase changes the picture. Once the effort stops, new notes go to the new app but a substantial amount of reference material is still in the old folder, so every lookup now begins with a decision about where to look. Time per retrieval has gone up, not down, and the increase is invisible per instance.
The correct answer here is neither continue nor remove. It is modify: finish the migration, or designate the old folder as archive only and accept that anything in it is read only history. Both remove the two places problem. Doing neither is the state most people stay in indefinitely.
The general shape of this recurs constantly. The tool was fine. The half finished transition was the problem, and no amount of evaluating the tool would have surfaced it.
Tools with nothing to maintain
Our tools have no accounts, no settings and no stored state, which means their maintenance cost is zero and their switching cost is a browser tab. That is a deliberate design choice, and it is the reason they survive the durability phase.
See the toolsWhere this test gives the wrong answer
It measures time and errors, so it will misjudge anything whose value is not primarily either.
Tools that are insurance. A backup service, a password manager or a version control system may save no time at all in a normal fortnight and still be clearly worth keeping, because their value is concentrated in a rare event. Do not run this test on those. The question there is what the bad day costs, not what the average day costs.
Tools that raise quality rather than speed. Something that makes work better while taking longer will fail a time based test and might still be right. Judge those on the output, and be honest that you are doing so.
Tools other people depend on. If a colleague relies on the shared visibility something provides, your personal net value is the wrong measure. This test is for your own workflow.
Anything you have used for under two weeks. Early friction is mostly learning, and learning ends. Adopt it properly first, then test.
One more honest caveat: fourteen days does not see monthly or quarterly work. If the tool exists mainly for a month end process, this test will not observe its best case, and you should extend the window rather than draw a conclusion from a fortnight that never contained the relevant task.
Frequently asked questions
How do I know whether a tool is worth keeping?
Compare recurring benefit against full recurring cost, including maintenance, switching and any new errors, after the novelty has worn off. If it survives four days of you not making any effort, it is probably real.
Why not just decide based on how it feels?
Because two things pull that judgement in the same direction. A new tool gets deliberate attention your old process had lost, and the setup time you invested makes you rate the result more kindly. Both inflate the same answer.
What if I skipped the baseline and already adopted the tool?
Run the baseline anyway by deliberately doing the task the old way for two or three instances. It feels wasteful and it is the only way to get a comparison you can trust. Failing that, run only the durability phase, which still answers the most important question.
Does this work for a whole suite rather than one tool?
Not well. Test one tool at a time, because with several changing at once you cannot attribute the result. If you have just adopted a suite, pick the component you are least sure about.
What counts as a maintenance cost?
Any time spent on the tool rather than on the work. Adjusting settings, fixing a sync, reorganising, tidying, updating a template, re-learning something after an interface change. If the activity would be pointless in the absence of the tool, it is maintenance.
The result was ambiguous. Now what?
Ambiguous almost always means modify. Change exactly one thing about how you use it, most often narrowing it to a single kind of work, and repeat the active week. Changing two things at once leaves you unable to say which one mattered.
Where to start this week
Pick the tool you are least certain about, which is usually the one you would feel slightly defensive about explaining. Spend three days measuring the task without it, one week using it deliberately, and four days not trying at all.
Then write two sentences: what it saves in a normal week, and what it costs in the same week. If you cannot write the second sentence, you have not been counting the recurring costs, and that is where the answer is hiding.
How this was put together: the test structure and the cost list come from our own experience adopting and abandoning tools while building the ones on this site. It is not research, there is no sample behind it, and we have quoted no percentages or average savings because we have measured none at a scale that would make such figures honest. The net weekly value calculation is arithmetic on your own numbers, so it holds regardless of ours. The three phase structure and the fourteen day window are judgement calls offered as a starting point, not findings.