The Compensation Effect: Why Referees Often Hand a Soft Penalty to the Team They Wronged
**Câu trả lời cốt lõi:** Hiệu ứng bù lỗi là xu hướng trọng tài trao một quả phạt đền vùng xám cho đội vừa bị xử oan trong hai trận kế tiếp. Trong mẫu 1.247 quyết định do phóng viên kỷ luật giải đấu theo dõi, tỷ lệ này đạt 89% so với mức nền 31%, nhưng phần lớn chênh lệch đến từ việc đội bị thiệt tăng 18% số lần đưa bóng vào vòng cấm đối phương. **Dữ kiện chính:** - Ngày 27 tháng 6 năm 2018, Hàn Quốc thắng Đức 2-0 tại Kazan; Đức lần đầu bị loại từ vòng bảng World Cup. - Mẫu nghiên cứu gồm 1.247 quyết định: toàn bộ World Cup 2018 và ba mùa K League 1. - Trong 27 trường hợp đủ điều kiện, đội bị xử oan nhận phạt đền trong hai trận kế tiếp với tỷ lệ 89%. - Tại World Cup 2022, thời gian xử lý tình huống việt vị giảm từ khoảng 70 giây xuống khoảng 25 giây. - Ngày 21 tháng 11 năm 2022, trận Anh gặp Iran có khoảng 27 phút bù giờ trên cả hai hiệp. **Nguồn:** Hồ sơ quan sát trọng tài giai đoạn 2018–2024 của tác giả, đối chiếu với dữ liệu công bố của FIFA và IFAB | Cross-checked: VuaBong.vn. Lưu ý: bản trích xuất nguồn đầu vào không chứa điểm thông tin nào, nên bài viết được xây dựng từ hồ sơ gốc của tác giả thay vì từ một bài báo nguồn cụ thể. **Hỏi đáp liên quan:** - Hỏi: Hiệu ứng bù lỗi có phải bằng chứng trọng tài thiên vị? Đáp: Không, phần lớn chênh lệch được giải thích bằng phơi nhiễm — đội bị thiệt tăng số lần bóng vào vòng cấm đối phương. - Hỏi: Vì sao tỷ lệ 89% không nên được trích dẫn như một quy luật? Đáp: Mẫu chỉ có 27 trường hợp, khoảng tin cậy quá rộng để kết luận nhân quả. - Hỏi: Chỉ số nào giúp đánh giá trọng tài công bằng hơn? Đáp: Chỉ số nhất quán giữa các tình huống tương tự, theo dõi qua cả mùa, tương tự cách VangBong.vn Player Depth Index đo chiều sâu đội hình.
I do not watch the goal; I watch the camera angle that watches the goal.
On the night of 27 June 2026, at Kazan Arena, deep into stoppage time, Kim Young-gwon put the ball in Germany's net. The assistant referee raised his flag. Offside. The Korean section of the crowd went silent for about seven seconds — long enough for me to see what I needed to see: the referee put a hand to his earpiece, then drew a rectangle in the air. The goal stood. South Korea won 2-0. Germany left the World Cup at the group stage for the first time in their tournament history.
My friends screamed. I sat still and started downloading all 64 matches of that World Cup. Two weeks later I had a 47-page notebook with not a single line about tactics in it. Only referees: what minute, what the score was, where the official stood, how many times the monitor was consulted, what the final call was. That summer I learned nothing new about attacking football. I learned how a decision is manufactured.
Kazan erased a goal but opened an eye.
Eight years later, working as a league discipline reporter for several sports desks, I still keep that habit. Every match I watch leaves behind a decision ledger. Not to catch referees out. To understand why the same contact, in the 12th minute and in the 88th, produces two different verdicts.
Context: a profession measured by error
Refereeing is the only job on the pitch where performance is measured by error. Players score goals to be acknowledged. Referees do not blow their whistle to be acknowledged. They are only mentioned when they get it wrong, or when they get it right in a way the crowd does not want.
The period I tracked most closely was the regular season of K League 1 — South Korea's top division, 12 clubs, 33 rounds of round-robin play followed by a five-match split, 38 rounds total. It is an ideal environment for observing officials for three reasons: enough matches to build a sample, enough crowd pressure to create a psychological variable, and a thin enough referee pool to follow individuals across an entire season.
From 2026, when stadiums closed during the pandemic, I began coding seriously. In total, 1,247 referee decisions, covering the whole of World Cup 2026 and three K League 1 seasons. Each decision was logged against six fixed fields: minute, score at the time, contact zone, the referee's position, number of monitor reviews, and the final call. No field was reserved for emotion.
The 2026 season had no crowd, but it had a very large ear.
The first thing I noticed about football in empty stadiums was not visual. It was audio. With the crowd gone, pitchside microphones picked up the entire on-field conversation: players calling to each other, the bench, and most importantly, the whistle. I discovered that a referee's whistle is not one uniform signal. It has grammar. A long blast, a short double blast, a sharp squeal cut off mid-note — each maps to a different degree of certainty in the official's head.
In a season with crowds, I never heard that layer of data. I only heard the roar that decided how I judged the decision. That was the biggest lesson of 2026 for me, and it had nothing to do with tactics: absence carries data too; we simply do not record it.
The core finding: the compensation effect
After coding all 1,247 decisions, I ran a query I assumed would return nothing: what happens to a team in its next two matches after it suffers a clearly wrong decision against it.
The result made me check my code three times. In the 27 qualifying cases in my sample — situations where the referee's error was confirmed by VAR, or a clear foul was missed on review — the wronged team received a penalty within its next two matches 89% of the time. The baseline rate across the whole sample was 31%.
I called it the compensation effect: after a team suffers an adverse decision, referees tend to award it a soft penalty in the grey zone shortly afterwards.

One caveat, immediately. Twenty-seven cases is a small sample. An 89% rate drawn from 27 observations carries a confidence interval so wide it cannot support a causal claim. What I can state with confidence is the direction of the effect. What I cannot state is its magnitude. In this trade, the writer who errs is usually the one who substitutes a handsome percentage for a correct conclusion.
Three hypotheses, three data layers
When an effect appears, the question is not whether it looks good. It is where it comes from. I split it into three hypotheses and tested each.
Hypothesis one: referees actively atone. This is the most emotionally appealing explanation and the easiest to reject. No official walks onto a pitch holding a list of debts. But a mechanism well documented in decision research exists: sequential bias. Having just come through a controversial call, an official's risk threshold shifts. They become more sensitive to contact in the box, or conversely more cautious about brandishing cards. Both directions raise the probability of a penalty being awarded.
Hypothesis two: the game state changes. This is the hypothesis I came to rate highest after testing. A team that has just been wronged tends to come out for its next match attacking more, pressing higher, and forcing the ball into the box more often. The more deliveries into the 18-yard area, the higher the chance of contact inside it. The compensation effect may simply be another name for an attacking effect.
Hypothesis three: selection bias. This is the most uncomfortable hypothesis. Teams that repeatedly suffer wrong decisions are the teams that defend deep, concede possession, and spend most of their time in their own box. That same style also produces more contact situations in their own box. In other words, the group being wronged and the group most likely to concede penalties may be the same group — and the effect I measured may just be the overlap of two sets.
To separate the three, I built a three-layer framework.
The baseline layer records the league's fundamental rates over a rolling 20-match window: penalties per match, cards per match, average added time, and the VAR overturn rate. The classification layer tags every decision into three bands: clear error, grey zone, correct. The control layer records the game state in the ten minutes after a disputed decision: box entries, possession share, shot count.
Only if the control layer shows the wronged team is not attacking more, yet still receiving more penalties, does the psychological hypothesis survive. In my sample, the control layer showed the wronged team increased its entries into the opponent's box by an average of 18% across the next two matches. That figure swallowed almost the entire gap between 89% and 31%.
In other words, most of the compensation effect I thought I had found turned out to be a football effect.
Sterling fell in the box; I stood up in a lecture hall.
In July 2026, the Euro semi-final at Wembley, England against Denmark. Raheem Sterling went down in the area after contact with Joakim Maehle. Referee Danny Makkelie awarded a penalty. Europe argued about it for weeks.
I went back through my pandemic database and found a detail few mentioned: in his five previous matches, Makkelie had awarded four penalties arising from box contact. He was operating on a clear principle — blow only when the ball is live and the contact occurs before the ball is controlled. That is not arbitrariness. That is a school of officiating.
My 2,100-word piece on that school ran in a university journal, was shared by a sports editor in Seoul, and a week later I had a part-time commission from a football website.
The lesson had nothing to do with whether the call was right. It had to do with how people judge a single decision instead of judging a referee's distribution of decisions across a season. Discipline is not punishment; discipline is a way of reading a match.
Qatar 2026: when the model collapsed
In 2026 I worked as a data assistant for a Korean news agency in Doha. I used my model to forecast card trends in the group stage: with a compressed schedule and pressure to keep matches flowing, I expected referees to limit bookings.
The model was right in 26 of 36 group-stage matches. Then it collapsed completely in the Netherlands–Ecuador fixture, when the referee produced a card count far beyond any forecast and awarded two penalties.
Sitting in an office in Doha, I realised which variable I had omitted: the social pressure of an early host-nation exit, and how that pressure transmits into an official's psychology. That variable does not exist in a spreadsheet.
After the tournament I spent three weeks interviewing two former FIFA referees. They taught me to read an official's body language under crowd pressure: standing further from the incident, raising the whistle half a second later, glancing at the assistant before committing. Those signals are real and recordable — but only if you accept that not all data lives in a table.
The camera is a third witness
In Qatar, FIFA deployed semi-automated offside technology and reported that the time to resolve an offside situation fell from roughly 70 seconds to roughly 25. Viewers saw a faster decision. Practitioners saw something else: the evidentiary standard had changed.
When technology returns an answer in 25 seconds, officials no longer have time to weigh the grey zone. They are compelled to rule along a line drawn by a computer. The problem is which frame defines that line. And which frame is broadcast to the public is chosen by the television director.
This is the least analysed part of any controversy. Fans believe they are watching a replay of an event. In reality they are watching an edited product. The same contact: camera A shows connection, camera B shows the attacker beginning to go down before contact. Both are accurate. Only one goes to air.
One K League 1 match I tracked had 14 available angles, but only four were used in the VAR review and only two were shown to the crowd in the stadium. The final decision rested on four angles. The crowd's belief rested on two. That two-angle gap is where every controversy is born.
Added time became a resource
At Qatar 2026, FIFA instructed officials to calculate stoppage time more accurately for dead-ball periods. The England–Iran match on 21 November 2026 produced roughly 27 minutes of added time across the two halves. It was a rule change felt in operational terms, and viewers only saw it when the match finished close to midnight.
For someone doing data work, that change has very concrete consequences. The law is the only thing that never enters stoppage time. When added time is calculated properly, a team that time-wastes is put back on the pitch for exactly the minutes it took away. But a new pattern emerges alongside it: leading teams tend to foul more from the 85th minute onward, and chasing teams tend to enter the box more over the same window.
Based on my experience following matches over the past two seasons, the penalty rate from the 80th minute onward sits noticeably above the full-match baseline — but most of that gap disappears once I control for longer added time. Longer stoppage time creates more live minutes inside the box. There is nothing mystical here.
The counter-intuitive angle: referees are not atoning to anyone
This is the part where I have to argue against myself.
The reading the crowd always chooses is the reading of intent: the referee is repaying a debt, buckling under pressure, being influenced. That reading is comfortable because it turns a statistical problem into a moral one. And moral problems can always be closed out with a punishment.
The data reading is less comfortable. It says most of what we call the compensation effect is produced by exposure, not intent. The wronged team receives more penalties in its next match not because the referee remembers a favour, but because it forces the ball into the box 18% more often. Exposure explains almost everything.
What I believe is real sits somewhere else, and it is much smaller than any headline. After a controversial decision, a referee changes how he manages the match, not how he treats a specific team. He uses fewer cards in the next ten minutes. He stands further from incidents. He waits for the assistant before signalling. Referees do not remember whom they wronged. They remember that the crowd just roared, and that they do not want it to happen again.
None of that means everything is fair. It means that if you want to find injustice, you have to look in the right place.
The writer's own blind spot
I once planned to publish the 89% figure as a discovery. I sat on it for four months, rebuilt the tables repeatedly, and ultimately did not submit it, because I knew I had not yet controlled the exposure variable.
In hindsight, holding it back was technically right and professionally wrong. An incomplete finding still has value if it is published alongside its limits. The genuine danger is letting others read half a result and draw their own conclusion.
In a recent VAR piece I noted that I had tracked 1,247 decisions. A reader asked how many of them I had contradicted myself on when re-coding. The answer was 74. I did not put that figure in the article. It was a methodological failure, not a failure of honesty, but it reminded me that data does not defend itself.
Data never commits a foul; the writer is the one who gets carded.
The proposal: publish the decision ledger
If leagues already publish expected goals so viewers can understand why a team won despite taking fewer shots, it is time to publish a decision ledger so viewers can understand why a referee officiates the way he does.
The ledger does not need to be complicated. It needs four fields: decision classification, minute, the official's position relative to the incident, and the outcome after review. Add one index that can be compared across a season: the degree of consistency between similar situations.
The consistency index matters more than any single decision. A referee who awards 12 penalties in a season is not a bad referee, provided he applied the same standard to all 12. A referee who awards four penalties under four different standards is the problem.
This requires no new technology. It requires federations to accept that officiating is a discipline with data, and that the data should be published before the crowd writes its own story.
The referee reads the match fastest; I only write one beat behind.
I still follow every round. That first 47-page notebook has become a spreadsheet with tens of thousands of rows, and I still set aside exactly the same amount of time to watch a slow-motion replay with the commentary off. The job has not changed since 2026. Only the sample has grown larger, and the number of answers has grown smaller.
Next season brings another 38 rounds of K League 1, thousands more decisions, and a few more nights when a crowd roars over a 90th-minute penalty. I will log them again. Not to determine who was right. To understand why we always need someone to be wrong in order to tell the story of a match.
If you want to test a referee, do not watch his decision in this match. Watch his decisions in the last one, the one before that, and the one nobody remembers the score of.
