From Sanskriti Khandelwal | Product & Market Analysis
Remote Work Monitoring Is Colliding With the Trust That Made It Work
On this page
A meta-analysis of 94 independent samples covering 23,461 workers found no evidence that electronic performance monitoring improves performance. It did find higher worker stress, and that finding held regardless of how the monitoring was designed. Remote work monitoring is being bought on a promise the research record does not support, and paid for in the one input distributed teams cannot restock quickly.
Key takeaways
- The productivity case for monitoring is absent from the evidence, not merely contested. Ravid and colleagues pooled 94 samples and 23,461 participants in Personnel Psychology and reported no evidence that monitoring improves worker performance.
- Monitoring reliably changes behaviour, in the direction you did not buy. Across two studies, monitored employees were substantially more likely to break rules, because monitoring reduced their felt responsibility for their own conduct.
- Trust shows up in the accounts as retention, not as sentiment. A randomised trial of 1,612 employees published in Nature cut attrition by roughly a third with no measurable effect on performance grades over the following two years.
- The rules are diverging by jurisdiction, fast. Europe banned inferring emotions at work from February 2025 and fined one employer 32 million euros for excessive activity tracking. US federal guidance on the same question was withdrawn in the same month.
What "remote work monitoring" now covers
Direct answer. Remote work monitoring is any system that records what a distributed worker does, then converts that record into a judgement about them. The published evidence does not show it raising output. It does show higher stress, lower felt responsibility and more rule-breaking. Outcome contracts and team-level telemetry get you the visibility without the same cost.
The category grew by accretion, which is why most buyers cannot describe what they own. A firm signs up for endpoint security, then adds a time tracker, then a collaboration analytics module, and ends up with a behavioural record nobody designed on purpose.
The three layers most firms are running
The first layer is presence. Idle timers, session length, status colour in a chat tool, screenshots at intervals. It answers whether somebody was at the keyboard, which is a question about attendance dressed as a question about work.
The second layer is activity. Application and URL histories, keystroke and mouse counts, document and ticket events. This layer produces most of the data volume and almost none of the insight, because activity correlates with typing rather than with value delivered.
The third layer is inference. A model reads the first two layers and returns a score: productivity index, engagement risk, attrition risk, sentiment on internal messages. This is the layer that changed in the last three years, and it is the layer with the thinnest evidence base underneath it.
What AI actually added
Machine learning did not make the underlying measurements better. Keystroke counts are the same signal they were in 2015. What models added is the willingness to draw a conclusion from that signal and to attach it to a named person automatically.
That is a governance change disguised as a tooling change. A dashboard showing 4 hours of active time asks a manager to interpret it. A score reading "engagement risk: high" has already interpreted it, and the manager now has to argue with a number. The same substitution is happening in how AI is being written into performance review criteria, and the failure mode is identical in both places.
The performance case is missing from the research record
This is the part most vendor material skips, so state it precisely.
What the meta-analysis actually measured
Ravid, White, Tomczak, Miles and Behrend pooled 94 independent samples and 23,461 participants and published the result in Personnel Psychology in 2023. Their conclusion on the central question is one sentence long. The results provide no evidence that electronic performance monitoring improves worker performance.
That is a null result, not a finding of harm to output. It matters because it inverts the burden of proof. A firm buying monitoring to raise productivity is making a claim that decades of accumulated study have not been able to detect.
The stress finding is the one that did reach significance. The presence of monitoring was associated with increased worker stress, and the authors report that this held regardless of the characteristics of the monitoring. Design choices moved worker attitudes. They did not remove the stress effect.
Survey work lands in the same place. The American Psychological Association's 2023 Work in America survey found that 56% of monitored workers report feeling tense or stressed at work, against 40% of workers who are not monitored. The gaps on emotional exhaustion, on withdrawal from colleagues and on motivation all run in the same direction.
A peer-reviewed Canadian study published in Socius in 2024 went further and traced the mechanism. The authors worked from a national sample of 3,508 working Canadians surveyed in September 2021. They found surveillance was indirectly associated with increased psychological distress through three channels: reduced job autonomy, increased job pressure and privacy violation. Job pressure alone accounted for more than half of the indirect effect.
The same paper is honest about a complication that most coverage drops. On job satisfaction specifically, the negative indirect effects were balanced by a positive direct effect, and the net association was not significant. Surveillance did not make people hate their jobs. It made the job heavier.
Monitoring does change behaviour, just not the behaviour you bought
The null result on performance is not the same as no effect at all. Something does happen when the watching starts. It is simply not the thing written on the purchase order.
The agency mechanism
Chase Thiel, Julena Bonner, John Bush, David Welsh and Niharika Garud ran two studies and reported the result in Harvard Business Review in June 2022. Monitored employees were substantially more likely to break rules, including taking unapproved breaks, disregarding instructions, damaging property, taking equipment and deliberately working slowly.
The mechanism the authors identified is the interesting part. Monitoring reduced employees' sense of agency and personal responsibility for their own conduct. Once the system is watching, the moral weight of the decision shifts onto the system, and behaviour the person would otherwise refuse becomes available to them.
My reading is that this is the most important finding in the literature for anyone running a distributed team, and vendor material almost never quotes it. It says the control mechanism erodes the internal control it was meant to supplement.
A ten dollar device beat the entire stack
The clearest public illustration is not academic. In May 2024, Wells Fargo dismissed more than a dozen staff in its wealth and investment management unit. The dismissals followed a review of allegations involving simulation of keyboard activity, according to disclosures filed with FINRA and reported by Bloomberg.
The devices used cost roughly ten dollars and are sold openly. They nudge a mouse on a timer so that a session never registers as idle. Every layer of presence measurement in the industry is defeated by a small piece of plastic that anyone can order.
Read carefully, this is not primarily a story about dishonest employees. It is about a trivially gameable measurement being treated as evidence, inside a regulated firm, to the point where gaming it became a terminable offence. The measurement created the offence. That is what Charles Goodhart warned about, half a century before anybody shipped a productivity score.
Trust is the input remote work actually runs on
The argument that measurement erodes trust usually stalls at this point, because trust sounds unfalsifiable. It is not. It has a proxy that appears in the accounts, and there is a randomised trial attached to it.
What the retention experiment showed
Nicholas Bloom and colleagues randomised 1,612 employees at Trip.com in Shanghai into five days in the office or three, then followed them for two years. The result, published in Nature in 2024, was a reduction in attrition of about one third, against a control-group base of 7.2%, with null equivalence tests showing no effect on performance grades.
Set the politics aside and read what the experiment isolates. Granting discretion over where work happened produced a large, measurable retention gain and cost nothing in measured output. Discretion was the active ingredient in that result.
Monitoring is the systematic withdrawal of exactly that discretion. It is reasonable to expect the retention effect to run in reverse when you withdraw it, and any firm deploying monitoring should be tracking precisely that. Almost none do, which is why the trust cost stays invisible in the business case.
Productivity paranoia is a management measurement failure
Microsoft's Work Trend Index put a number on the underlying anxiety in September 2022. 85% of leaders said the shift to hybrid work made it hard to be confident that employees were being productive, while 87% of employees reported that they were productive.
That gap is not a factual dispute about output. It is a confidence problem inside management, and monitoring software is sold as its cure. The honest diagnosis is that managers lost the informal signals of an office and were never given a replacement definition of done. The same vacuum is what makes the middle management layer so exposed as agents absorb coordination work.
Buying surveillance to fix a definition problem is a category error, and it is an expensive one.
The adoption statistic everyone quotes cannot be checked
Search for how many employers monitor remote staff and you will find a tidy figure inside a minute. 78%, or 74%, or 80%, or 70% of large employers. The numbers are quoted with total confidence and they appear in dozens of places.
I went looking for the underlying instrument behind each of them and could not reach a single one. The trail runs through monitoring software vendors' own blogs, statistics roundups and aggregators that cite each other. Six restatements of a figure is one source, and in several cases the original survey has no published sample size, no fielding window and no questionnaire.
So this post publishes no adoption percentage at all. That choice costs the piece a quotable headline number. It is the correct trade, because a prevalence figure with no methodology behind it is exactly the kind of claim that would undermine everything else here. The same discipline applies to the confident numbers circulating about AI's effect on aggregate productivity, where the measured record is far thinner than the commentary suggests.
Monitoring capability now ships inside tools that firms buy for other reasons, which means adoption is frequently a default rather than a decision. If you want to know your own number, read your admin consoles rather than a chart.
Two legal floors, and they are moving apart
Anyone running a distributed team across borders is now managing two incompatible trajectories at the same time.
Europe drew two hard lines
The first line is data protection. On 27 December 2023 the CNIL fined Amazon France Logistique 32 million euros for an excessively intrusive system that tracked warehouse activity through handheld scanners, and for retaining the resulting indicators for 31 days. The finding was not that tracking is illegal. It was that this volume of tracking failed data minimisation.
The second line is the AI Act. Since 2 February 2025, Article 5(1)(f) of Regulation (EU) 2024/1689 has prohibited AI systems that infer emotions of a person in the workplace from biometric data, with narrow medical and safety exceptions. Voice, face, gait and keystroke rhythm all fall in scope where the system reads a body and returns an emotional conclusion. Penalties reach 35 million euros or 7% of global turnover. If you are mapping exposure, start with the EU AI Act transparency obligations checklist.
The US federal position moved the other way
In October 2022 the NLRB General Counsel issued memorandum GC 23-02. It urged the Board to treat pervasive electronic surveillance and algorithmic management as presumptively unlawful where that would tend to interfere with protected activity. On 14 February 2025, Acting General Counsel William Cowen rescinded that memorandum along with 28 others.
State law then failed to fill the gap in the largest market. California's SB 7, the No Robo Bosses Act, would have required notice before automated decision systems affected employment and barred sole reliance on them for discipline or termination. Governor Newsom vetoed it on 13 October 2025, calling the notification requirements unfocused. A revised bill is expected in the next session.
The practical consequence for a distributed team is that your monitoring configuration is now a jurisdiction question rather than a company policy question. A single global default is very likely unlawful somewhere in your footprint. Recording tools sit in the same trap, which is why meeting note-taker consent and training rights have become a procurement issue rather than an IT setting.
Four designs that buy visibility at a lower trust cost
The alternative to monitoring is not blind faith in everybody. It is measuring the thing you actually care about, at the level where measurement does the least damage.
| Design | What you measure | Where it fails |
|---|---|---|
| Outcome contracts | Agreed deliverables, dates and quality bar, set before the period starts | Weak where work is genuinely interrupt-driven and cannot be scheduled |
| Team-level telemetry | Cycle time, flow efficiency and rework, aggregated so no individual is scored | Small teams make aggregation thin, and re-identification becomes possible |
| Declared availability | Hours a person commits to being reachable, self-declared and published | Relies on honesty, which is the point of it and also the objection to it |
| Exception review | Only cases that miss the agreed bar, reviewed by a human with context | Catches problems late, so it needs short delivery cycles to work at all |
These four are not equally evidenced. The autonomy pathway behind them is supported by the Nature trial and the Socius mediation analysis. The specific combination above is an editorial recommendation, not a measured result.
Measure at the smallest unit that answers the question, and no smaller than that. Delivery predictability is a team property, so measure the team. Individual keystroke data answers no question a manager actually has, and it carries every cost the research documents.
Two operational rails follow from that rule. Tell people what is collected before it is switched on, in specific terms, because the meta-analysis found transparency improves attitudes even though it does not remove the stress effect. And publish the counter-metric alongside the productivity metric, which means voluntary attrition in monitored teams against comparable unmonitored ones. If the tool is working, that comparison will survive being looked at. The developer population is where this shows up first, and the widening distance between stated trust in AI tooling and actual usage is the same signal in a different form.
Where this argument is weakest
Four things cut against the case above, and two of them are serious.
A null result in a meta-analysis is not proof of no effect. It can also reflect heterogeneous studies, weak measures of performance and publication practices in the underlying literature. The pooled samples are dominated by pre-2020 work, and very little of it tested AI-scored monitoring of the kind being sold now.
The APA and Socius findings are also cross-sectional and self-reported. Stressed workers may be more likely to notice and to report monitoring, which would inflate the gap between the two groups. Neither dataset can settle direction of causation on its own, and I would not present them as though they could.
There is also a legitimate case for monitoring that this post has mostly set aside. Regulated trading, healthcare records, payment systems and safety-critical operations carry statutory record-keeping duties, and in those contexts monitoring is not a productivity theory but a legal obligation. The mistake firms make is extending the compliance stack, built for a narrow duty, across a whole workforce that has no such duty attached to it.
Finally, the strongest version of the opposing argument deserves stating properly. Trust does not scale automatically, and a large distributed organisation cannot run on personal relationships alone. Some measurement is not a substitute for trust but a precondition for extending it to people you have never met. The question worth arguing about is which measurement, at what level of granularity, disclosed how, and that is a design question rather than a moral one. It is also connected to pace, because sustained delivery pressure produces the same withdrawal symptoms as surveillance does, as the pattern behind engineering burnout under AI-accelerated delivery cycles shows.
Frequently asked questions
Does employee monitoring actually increase productivity?
There is no published evidence that it does. The largest meta-analysis on the question pooled 94 independent samples and 23,461 participants in Personnel Psychology in 2023, and concluded that the results provide no evidence that electronic performance monitoring improves worker performance. The same analysis found monitoring was associated with increased worker stress regardless of how it was configured, so the measurable effects run against the buyer's stated goal.
Is AI employee monitoring legal in the EU?
Some of it is not. Since 2 February 2025, Article 5(1)(f) of the EU AI Act has prohibited AI systems that infer emotions of a person in the workplace from biometric data, with narrow medical and safety exceptions. Penalties reach 35 million euros or 7% of global annual turnover. Activity monitoring that does not infer emotion remains lawful in principle, but must still satisfy GDPR data minimisation, as the CNIL's 32 million euro fine against Amazon France Logistique demonstrated.
Can my employer monitor me while I work from home in the US?
Generally yes, on employer-provided equipment, subject to state notice laws such as New York's and Connecticut's. The federal position loosened in 2025. NLRB memorandum GC 23-02, which had urged treating pervasive surveillance as presumptively unlawful where it chills protected activity, was rescinded on 14 February 2025. California's No Robo Bosses Act, which would have added notice duties, was vetoed in October 2025.
Why do employees use mouse jigglers?
Because presence measurement rewards the appearance of activity rather than the delivery of work. Devices that nudge a mouse on a timer cost around ten dollars and defeat idle detection entirely. In May 2024, Wells Fargo dismissed more than a dozen staff after reviewing allegations involving simulation of keyboard activity, according to FINRA disclosures reported by Bloomberg. The episode shows a gameable metric being treated as evidence about a person.
What should companies measure instead of activity for remote teams?
Measure at the smallest unit that answers your actual question. For delivery predictability that is the team, using cycle time, flow efficiency and rework rates aggregated so no individual is scored. Pair it with outcome contracts agreed before the period starts, and review exceptions rather than everything. Then publish voluntary attrition alongside the productivity number, because that is where the cost of getting it wrong appears first.
Does transparent monitoring reduce the harm?
It helps with attitudes and does not remove the stress effect. The 2023 meta-analysis found that organisations monitoring more transparently and less invasively can expect more positive worker attitudes, while the association between monitoring and stress held regardless of monitoring characteristics. Disclosure is therefore necessary and insufficient. It is the minimum standard rather than a mitigation that lets you collect more.
Where to start this week
Open your admin consoles before you open a vendor deck. Most firms are running more collection than anyone consciously chose. Write down every signal currently collected, who can see it and how long it is retained. That inventory takes an afternoon and it frequently ends the conversation on its own.
Then run one comparison you can defend in a room. Take voluntary attrition over the last four quarters in your most monitored team and your least monitored team of similar size and function. Two numbers, one page, no software purchase required. If the monitored team is losing more people, you have measured the trust cost in the only currency a finance function accepts.
Related on measurement
The same design question shows up in AI-written performance review criteria, and the governance version of it sits in the EU AI Act transparency checklist.
References
- Ravid, White, Tomczak, Miles and Behrend, A meta-analysis of the effects of electronic performance monitoring on work outcomes, Personnel Psychology 76(1), 2023. K = 94, N = 23,461. Used for the performance, stress and transparency findings.
- Thiel, Bonner, Bush, Welsh and Garud, Monitoring Employees Makes Them More Likely to Break Rules, Harvard Business Review, 27 June 2022. Used for the rule-breaking result and the agency mechanism.
- American Psychological Association, Electronically monitoring your employees? It's impacting their mental health, 2023 Work in America survey. Used for all monitored against unmonitored percentages.
- Bloom, Han and Liang, Hybrid working from home improves retention without damaging performance, Nature, June 2024. Used for the 1,612-employee trial, the attrition reduction and the performance null result.
- Glavin, Bierman and Schieman, Private Eyes, They See Your Every Move, Socius, 2024. Used for the 3,508-worker sample and the three distress pathways.
- Herbert Smith Freehills Kramer, French regulator's 32 million euro fine against Amazon France Logistique, February 2024. Used for the CNIL decision and its grounds.
- National Labor Relations Board, GC 25-05, Rescission of Certain General Counsel Memoranda, 14 February 2025. Used for the withdrawal of GC 23-02.
- Bloomberg, Wells Fargo Fires Over a Dozen for Simulation of Keyboard Activity, 13 June 2024. Used for the FINRA disclosure detail.
Weakest thing about this source base: the APA page would not render for direct reading during research. Its percentages were taken from three consistent retrievals of the same 2023 survey rather than from the report itself, and should be re-checked against the source document. The EU AI Act prohibition and the California veto are cited to the legal instruments without a linked secondary source. Regulation (EU) 2024/1689 Article 5(1)(f) applied from 2 February 2025, and California SB 7 was vetoed on 13 October 2025.
Related reading