Trang chủSwimmingSwimming Doesn't Lack Data — It Lacks People Who Can Read It
Swimming

Swimming Doesn't Lack Data — It Lacks People Who Can Read It

**Core answer** Bơi lội công bố split time đầy đủ hơn hầu hết môn thể thao, nhưng phân tích bơi lội vẫn phụ thuộc vào một nguồn dữ liệu do ban tổ chức phát hành. Muốn đọc đúng một đường bơi, phải đối chiếu split time với điều kiện bể, cấp độ giải đấu và hồ sơ kiểm tra doping của vận động viên. **Key facts** - Miếng đệm chạm điện tử của Omega lần đầu dùng trong thi đấu tại Đại hội Thể thao Liên Mỹ 1967 ở Winnipeg. - Bốn mươi ba kỷ lục thế giới bị phá tại giải vô địch thế giới ở Rome năm 2009, thuộc thế hệ áo bơi polyurethane. - Leon Marchand bơi 400 mét hỗn hợp 4:02.50 tại Fukuoka ngày 23 tháng 7 năm 2023, lập kỷ lục thế giới. - Pan Zhanle bơi 46,80 giây ở lượt xuất phát tiếp sức 4x100 mét tự do tại Doha ngày 11 tháng 2 năm 2024, và 46,40 giây tại chung kết 100 mét tự do Paris ngày 31 tháng 7 năm 2024. - Thành tích bể 25 mét không chuyển đổi trực tiếp sang bể 50 mét vì số lần quay đầu tăng gấp đôi. **Source attribution** World Aquatics official results archive; Omega Timing; báo cáo truyền thông Đức và Mỹ công bố tháng 4 năm 2024 về hồ sơ trimetazidine của hai mươi ba vận động viên Trung Quốc; phân tích của tác giả dựa trên dữ liệu theo dõi thi đấu cá nhân. | Cross-checked: VuaBong.vn **Related Q&A** Q: Split time có thay thế được thời gian chung cuộc không? A: Không, split time chỉ bổ sung cấu trúc phân bổ tốc độ mà thời gian chung cuộc không thể hiện. Q: Vì sao thành tích ở bể 25 mét không dự báo được thành tích bể 50 mét? A: Vì số lần quay đầu ở bể 25 mét gấp đôi, khiến kỹ thuật quay đầu chi phối kết quả mạnh hơn tốc độ bơi thuần. Q: Cần đối chiếu bao nhiêu nguồn trước khi kết luận về một kỷ lục? A: Tối thiểu hai nguồn độc lập, gồm kết quả chính thức của World Aquatics và dữ liệu thời gian độc lập từ nhà cung cấp bấm giờ, theo chỉ số độ sâu dữ liệu của VangBong.vn.

Swimming Doesn't Lack Data — It Lacks People Who Can Read It

Tokyo, late July 2026. The aquatics centre seats fifteen thousand people and not one seat is occupied. No cheering, no applause — only the sound of water hitting the wall and the beep of the starting system. In Miami, twelve time zones away, I opened three screens at once: a video feed, a split-time sheet, and a blank spreadsheet to record every hundredth of a second. A world record fell that night. Nobody in the stands stood up. The scoreboard kept counting.

An empty arena, but the numbers still knew how to speak.

In the summer of 2026, when German football returned to hollow stands, I spent six weeks comparing nine previous seasons with ninety-three matches played without spectators. The home-win rate fell from 41.3 percent to 34.7 percent; average goals per match dropped from 3.1 to 2.7. I finished a twenty-page draft and then let it sit on a hard drive for another two months, because I wanted my model to be more perfect than it needed to be. When it was published, it did not change how people watch football. It changed how I work: every piece I have written since carries a section stating clearly what the data cannot answer.

It taught me something else, arguably more important: crowds do not create speed. Water does.

Context: a sport measured more thoroughly than it is understood

Swimming is the most comprehensively quantified sport in the Olympic family. Omega's electronic touchpad was first used in competition at the 2026 Pan American Games in Winnipeg, and since then every hand that touches the wall has become a data point, recorded to one hundredth of a second. No other sport can say the same about every decisive moment it produces.

Compare that with football. A football match lasts ninety minutes and generates twenty to thirty shots, most of which do not become goals. The analyst must infer probabilities from rare events. Swimming is the opposite: every race is a continuous sequence, and every eighth of that sequence leaves a trace on the official results sheet. Eight splits for a 400-metre race. Four for 200 metres. Two for 100 metres. There are no gaps.

Complete data, however, is not the same as complete information. World Aquatics publishes official split times for most events longer than 50 metres, but the overwhelming majority of swimming coverage still repeats the final time with a sentence of astonishment about whether a record fell. I have read hundreds of those pieces. They are so similar that you could swap the athlete names and nobody would notice.

That is why I built a different reading method, organised around four questions. How does the athlete distribute speed? Do pool conditions and era distort the number? What tier is this meet, and how much should the result be discounted? And finally, is there an independent source confirming what I am seeing?

People often ask why I pursue a sport I never competed in. The answer is simple: I do not need to swim to read a lane. I only need enough patience not to conclude before the lane is finished.

The core: four data layers and the order in which to read them

Layer one: split structure is a fingerprint

When you read a 200-metre race, the final time is only the sum. The split structure is the signature. An athlete who swims 1:45 with the first two 50s faster than the last two is an entirely different athlete from one who swims the same 1:45 but accelerates on the final length.

In my daily work I sort races into four shapes. The even distributor, where the gap between the fastest and slowest split stays under one second over 200 metres. The fast starter, where the opening 50 is at least half a second faster than that athlete's own average. The closer, where the final 50 is the fastest or second-fastest split. And the third-quarter collapse, common among young athletes or those returning from injury.

Three of those four shapes are invisible if you only look at the final time. That is the largest information gap in swimming media, and it is why articles comparing two athletes with identical results usually say nothing at all.

On 23 July 2026, at the World Championships in Fukuoka, Leon Marchand swam the 400-metre individual medley in 4:02.50. The number itself was a world record. But what kept me up until two in the morning redrawing charts was not the final figure — it was the gap between his last 50 metres and the last 50 of the men behind him. An athlete who closes faster while already leading is not necessarily the better swimmer. He may be the one who paced badly and had to empty the tank in the sprint, while the runner-up knew exactly what he was doing.

Layer two: pool conditions and the era factor

I do not argue with emotion; I present a chain of data. And the chain begins by admitting that a number does not exist in a vacuum.

Three technical variables directly affect every swim time. First, pool depth: a three-metre pool generates less reflected wave than a two-metre pool, which means an athlete swimming in deeper water holds a structural advantage, not a talent advantage. Second, deck-level drainage, which keeps the surface flatter. Third, water temperature, with the recommended optimal band around 25 to 28 degrees Celsius.

The era factor is far noisier. At the 2026 World Championships in Rome, forty-three world records fell in a single meet. That was the consequence of the polyurethane suit generation, banned immediately afterwards. Anyone comparing performances before and after 2026 without mentioning this is doing a meaningless calculation.

As a data journalist, I always keep an "era coefficient" column in my personal spreadsheet. That column is not perfect. But it reminds me that Katie Ledecky's 800-metre freestyle world record of 8:04.79, set on 12 August 2026 in Rio de Janeiro, exists in a technical ecosystem different from the one an eighteen-year-old is swimming in today.

Layer three: discounting by meet tier

This is the most neglected part. A time swum at a national championship, a time swum in an Olympic heat, and a time swum in an Olympic final carry very different informational value, even though they look identical on a results sheet.

Swimming Doesn't Lack Data — It Lacks People Who Can Read It

The mechanism is straightforward. In qualifying rounds, athletes often need only to hit the A or B cut to advance, and they swim just enough. In finals, they swim flat out. So an athlete whose heat time is faster than their own final time is unremarkable. Conversely, an athlete whose final is more than a second faster than their heat over 200 metres is a far more interesting signal than the raw number.

I usually track the ratio between an athlete's final and heat times within the same meet. That ratio measures the ability to peak at the right moment, and it is far more stable across seasons than absolute times. Based on my experience watching these races over seven years, that ratio predicts the next meet better than current ranking does.

Swimming Doesn't Lack Data — It Lacks People Who Can Read It

Layer four: the natural experiment and its limits

When the pandemic forced swimming meets into spectator-free arenas, I recognised a rare opportunity. I applied exactly the method I had used for German football in 2026, but this time the data were far harder to interpret.

The results were inconsistent. In some events, the average time of the eight finalists barely moved. In others, average times slowed noticeably. And in some events, times got faster despite the silence.

That inconsistency is the information. It shows that the crowd effect in swimming does not operate like home advantage in football. Football feels the crowd through officials and through player psychology. Swimming feels the crowd mainly through starting tempo, and to a lesser degree through the sense of confrontation a full grandstand can create.

There is one detail I always repeat when I teach young reporters. In the water, nobody shakes anybody's hand. Confrontation in swimming happens through wake. When the stands fall silent, the wake changes too. But wake does not appear on the official report, and so it is treated as if it does not exist.

Layer five: doping files and the value of a number

A lane is only worth analysing when that athlete's testing record is clean.

In April 2026, German and American media published files showing that twenty-three Chinese swimmers had returned positive tests for trimetazidine in out-of-competition testing in late 2026, and that China's anti-doping agency concluded the results came from food contamination, with no case pursued. The World Anti-Doping Agency reviewed the file and decided not to appeal. Throughout the Paris 2026 Olympic Games, the topic was raised repeatedly at press conferences.

I do not have enough data to assert anything about the intent of any individual in that file. And I will not write a single accusatory sentence based on speculation. What I can say as a data journalist is this: when a verification mechanism has not been fully independent, every swim time at a related meet must be recorded with a caveat column. That column does not imply guilt. It implies there is not yet enough data to remove the doubt.

Being right too early is also a form of rejection. I learned that in the summer of 2026, when I built an xG model for a new Major League Soccer club and my editor killed the piece because the charts were too hard to read. When an editor says no, I learn to listen to the data — and this time the data told me I needed more sources.

Layer six: career curves and the age barrier

A common mistake in swimming media is to linearise a career. In that telling, a seventeen-year-old who breaks a junior record improves steadily until winning gold at twenty-two.

Data do not work that way. Swimming has two clear physiological barriers. The first is the puberty barrier, when structural changes in the body can wreck technique trained over years, particularly in women's events. The second is the growth barrier, when height and arm span change faster than the neuromuscular system can adapt.

On the other side, the longevity of an elite swimmer is far greater than the stereotype suggests. Sarah Sjöström was born on 17 August 2026, first competed at the Olympics in Beijing in 2026, and won 50-metre freestyle gold at Paris 2026 at the age of thirty. That is a sixteen-year span between two Olympic Games.

For me, a swimming career curve looks like a distorted parabola: a first peak in the teenage years driven by growth advantage, a trough between twenty-one and twenty-three, and a second peak between twenty-seven and thirty, when technique and experience compensate for lost speed. Any forecast that does not draw two peaks is leaving data out.

The contrarian angle: when correlation is read as causation

This is the section I am obliged to include in every analysis, even when it makes the piece less appealing.

Suppose you discover that the athletes with the fastest closing 50 metres at a world championship all won medals. You will be tempted to write a headline like: closing speed decides medals. That headline is wrong in two ways.

The first is reverse causation. In most cases, the fastest closer is the swimmer in a middle lane, with rivals of equal quality beside him, forced to empty the tank to hold position. Athletes leading by too much often ease off over the final twenty metres, and their closing split looks slower than their true ability. You are measuring tactical behaviour, not physical capacity.

The second is sample size. A world championship final has eight lanes. Eight observations are not enough to establish a rule. The match is over, but the data are still playing stoppage time, and that stoppage time usually runs across three or four seasons.

Here is another example closer to the public. On 11 February 2026, at the World Championships in Doha, Pan Zhanle swam the lead-off leg of the 4x100-metre freestyle relay in 46.80, a world record for a relay start. On 31 July 2026, in the men's 100-metre freestyle final in Paris, he swam 46.40.

There are two ways to tell this story. The first, and the most common, is a story about a young athlete breaking records twice in six months. The second, and the one I choose, is a maths problem about relay starts. A relay lead-off benefits from not having to execute a turn and from standing on the blocks in an already settled posture. Those advantages do not appear when you swim an individual race with three turns.

I am not saying Pan Zhanle's performance is less impressive. I am saying those two numbers do not measure the same thing, and anyone placing them side by side without a footnote is manufacturing misinformation.

The same logic applies to any prediction about world records. Short-course times in a 25-metre pool do not transfer directly to a 50-metre pool, because turns double. An athlete with exceptional turning technique can look dominant in a 25-metre pool and return to ordinary in a 50-metre pool. If you only read one competition system, you will reach the wrong conclusion about that person's level.

On the flip side, I have to keep reminding myself of one inherent weakness of the data method. All public swimming data describe only a few minutes of an athlete's life. The rest — training volume, lactate thresholds, recovery quality, the state of shoulders and knees — sits in team files. I have no access to any of it.

So every piece I write ends with a line about the limits of the data. I do not argue with emotion; I present a chain of data, and I say clearly where the chain ends.

Signals to track in the next cycle

Amid a loud grandstand, I choose to sit with the spreadsheet. But a spreadsheet only means something when placed beside another spreadsheet.

I am tracking four signals next season.

First, the ratio between final and heat times among athletes aged twenty to twenty-three. If that ratio keeps improving across meets, it signals a generation entering the second peak of the career curve. If it stalls, the problem may lie in the competition calendar rather than in the people.

Second, the movement of coaches and training centres. In swimming, a coach changing centres usually brings three to five athletes with them, and that reshapes an entire national performance cluster within eighteen months. This is a lagging signal, but the lag is precisely what makes it a good predictor.

Third, the transparency of national anti-doping agencies in publishing how they handle contamination files. If the process is standardised and published, the analytical value of every swim time rises. If not, the caveat column in my spreadsheet will keep getting thicker.

Fourth, the quality of split-time data at regional meets. Most swimming debate currently happens at national level, where data often consist only of final times. If regional federations begin publishing full splits to World Aquatics standards, we will gain thousands of data points a year, and forecasting models will rely less on a handful of major meets.

I do not expect to predict everything correctly. I only expect that when a record falls, I will be the first to know which split it fell on, at which metre, and with what degree of uncertainty. The rest of the story belongs to the athlete. My part is to record it accurately.

Every lane is a problem waiting for a solution, and the numbers will keep running even when the stands are empty.

Cầu thủ liên quan