← Back to the latest articles
QUESTION 20Sports, Gender & Fairness

Does One Bad Call Become Evidence Against All Women?

A female umpire at Koshien, gender equality, ability differences, and the proper subject of criticism

Gender equality does not require us to insist that men and women have no average differences whatsoever. It is a commitment not to turn a group average—or one person's failure—into a verdict on the individual before us.

On August 7, 2026, a call by second-base umpire Kana Sato drew intense attention during Ariake's game against Ritsumeikan Uji at the 108th National High School Baseball Championship.

In the sixth inning, a Ritsumeikan Uji runner was initially called out at second. Because the fielder had not completed the catch, the umpiring crew conferred and changed the call to safe. After the game, Sato explained that she had realized her mistake and herself asked the other umpires to discuss changing the call. In the seventh, Ariake challenged an out call on a steal of second. The tournament's first video review was conducted, but that call stood.

According to the publicly reported sequence, then, the sixth inning included a missed call that Sato herself acknowledged. But the seventh-inning decision cannot simply be declared a “second bad call”: it remained out after video review.

Two reactions collided on social media.

One side said that women cannot umpire boys' high-school baseball, or that equality had been placed above competence. The other said that criticizing a female umpire was itself sexist, or that this historic first should simply be celebrated.

Are those really the only choices?

If a call was wrong, criticism and review are necessary. For high-school players, Koshien is a place where three years of effort can be affected by a single decision. Yet one mistake by one woman cannot settle the fitness of women as a whole.

The central claim of this essay is:

Gender equality does not exempt women from criticism. It returns the subject of criticism from “women” to “the call.”

Going one step further, equality is not the claim that men and women are identical in every respect. It is a social commitment not to use an average about a group as a verdict on the individual standing before us.

1. First, separate the event into two calls

Before debating the meaning of the incident, we need to separate the facts.

The sixth inning—error and correction

With no outs and a runner on first in the bottom of the sixth, a sacrifice bunt was fielded and thrown from third to second. The player covering second dropped the ball without completing the catch, but the runner was initially called out. The umpires then conferred and changed the call to safe because there had been no complete catch.

Sato herself said that the decision had been her mistake and that she asked the other umpires to change it. There is therefore a factual basis for saying that the sixth-inning call was wrong.

But another fact belongs beside it. Rather than defend an announced call by force of authority, she asked for a conference and tried to return the correct ruling to the players.

Correcting an error does not erase it. Yet admitting an error is not worthless either. Professional responsibility includes not only trying to be right the first time, but also being corrigible when one discovers that one was wrong.

The seventh inning—doubt and review

In the seventh, an Ariake runner attempted to steal second and was called out. Ariake requested video review, but the decision was not changed.

Spectators may say that the runner looked safe. But it is careless to disregard the official review and count this as a missed call with the same confidence as the corrected sixth-inning decision. A questionable call and an acknowledged, corrected error are not the same thing.

The distinction matters. We need not soften facts in order to oppose sexism, but neither should we inflate the number of errors in order to attack a woman.

2. Criticizing a call and explaining it by gender are different acts

A criticism of the play could read:

On the sixth-inning play, the runner was called out without confirming a completed catch. Video should be used to examine the umpire's positioning, line of sight, and why the decision was made too quickly.

This criticism concerns a particular act, a rule, evidence, and a path to improvement. It would make sense whatever the umpire's gender.

Now compare it with:

This proves that women cannot umpire boys' high-school baseball.

That sentence uses one decision to make a claim not about the call but about the nature of women as a group. Its problem is not merely that it is rude. The inference does not hold.

It contains at least four leaps:

  1. One call was wrong.
  2. This umpire is incompetent.
  3. Female umpires are incompetent.
  4. Women should not be appointed as umpires.

The first claim can be tested against video and the rules. The second requires a body of evidence about accuracy over time, positioning, rules knowledge, and game management. The third requires evidence comparing many male and female umpires under comparable conditions. The fourth is an ethical and policy judgment involving rights, opportunity, and institutional design.

A single video can directly establish, at most, the first step. On social media, however, a few seconds of footage can be made to jump all the way to the fourth.

I will call this the gendering of a bad call.

To gender a bad call is to explain it through the umpire's sex before examining technique, positioning, experience, attention, pressure, or crew coordination. At that moment, criticism stops being information for technical improvement and becomes evidence used to confirm an image of a group that the critic already held.

3. Defending women must not mean pretending the mistake did not happen

Should criticism of the umpire therefore be restrained? No.

To say that a call must not be questioned because she was among the first female umpires at the tournament would also fail to treat her as an equal professional. It would demand accuracy from male umpires while asking only symbolic value from women.

Low expectations offered in kindness can sometimes reach the same destination as open hostility. Both refuse to see a woman as an individual decision-maker and instead treat her as a symbol marked “woman.”

  • Hostile prejudice says, “She failed because she is a woman.”
  • Overprotection says, “Her failure should not be questioned because she is a woman.”
  • Equal evaluation says, “This call was wrong. Independently of gender, let us examine its cause and how to improve.”

Equality does not mean unconditional praise. It means the same standards, the same accountability, the same opportunities to learn, and the same right to try again.

When a male umpire misses a call, that one error does not exclude men as a group from umpiring. However severe the assessment, it is ordinarily treated as his call. Women likewise need both the right to fail and the right to recover from failure.

4. Why does the failure of a “first woman” become evidence about a group?

In an occupation containing many women, one woman's failure is more easily seen as one person's failure. Where women are rare, one woman is made to represent “women.”

Sociologist Rosabeth Moss Kanter's theory of tokenism described how numerical minorities become highly visible and are more readily perceived as symbols of a category than as individuals. More recent research likewise reports that tokenized people face intense scrutiny, isolation, and pressure to represent their group.

I will call this burden the representative tax.

The representative tax is paid when a member of a minority is made responsible not only for the result of her own work, but also for the reputation of people with the same trait whom she has never met.

If male umpire A errs, people say, “A was wrong.” If female umpire B errs, they say, “Women really cannot umpire.” B then pays twice for one mistake. She loses standing as a professional and is also charged with damaging the prospects of her group.

The tax is also asymmetrical:

  • A minority member's failure is generalized into a property of the group.
  • A minority member's success is individualized as an exception.

If a woman can make one hundred correct calls and be described merely as talented, yet make one error and be taken as proof that women are unfit, success and failure are not being counted in the same way. That is not meritocracy. It is a double standard built into the accounting of evidence.

5. What history produced the idea of gender equality?

Gender equality did not begin because science had proved that there were no differences between men and women. It began from a more fundamental question:

May a person's role in life be fixed in advance solely by an attribute of birth?

When “nature” became a reason for exclusion

For much of history, women were excluded from politics, education, scholarship, professions, and sports in the name of what was supposedly natural. They were described as emotional, passive, physically weak, and unsuited to authority.

But the argument contained a circle:

  1. Assume women lack ability and deny them education and experience.
  2. Treat the small number of experienced women as evidence that women lack ability.
  3. Use that “evidence” to deny still more opportunities.

This is less an observation of nature than an institution using an outcome it produced as proof of its own correctness.

Wollstonecraft—ask about education before ability

In A Vindication of the Rights of Woman (1792), Mary Wollstonecraft argued that if women appeared less rational, this might be the result of a society that had not given them an education capable of cultivating reason.

Her important point was not to stipulate that every female ability must be identical to every male ability. It was that a fair comparison of ability first requires fair conditions for developing it.

Applied to umpiring, we cannot deny a group access to national clinics, live games, video review, mentoring by senior umpires, and assignments to important contests—and then infer from its shortage of experienced officials that its members are unsuited to the work.

J. S. Mill—an unequal society cannot measure “women's nature”

In The Subjection of Women (1869), John Stuart Mill argued that much of what his age called “women's nature” may have been artificially produced by subordination and biased education.

In contemporary terms, Mill was objecting that the conditions of comparison had not been controlled. If two groups have long been given different education, expectations, and roles, we cannot simply attribute every resulting difference to inborn nature.

He treated society as a kind of experiment. Only when people are free to try can we discover who is suited to what. Closing the gate in advance and then citing the absence of women's achievements inside it is to refuse the experiment while announcing its conclusion.

Beauvoir—men appear as “human,” women as “women”

In The Second Sex, Simone de Beauvoir analyzed a structure in which man is treated as the unmarked human standard and woman as the sex-marked “Other.”

Baseball broadcasts rarely say “male umpire,” yet repeatedly say “female umpire.” Men are umpires; women become umpires-who-are-women.

When gender is always placed first, calls are interpreted through it. A man's missed call is human error, while a woman's is more readily treated as an expression of womanhood.

Elizabeth Anderson—equality is an equal relation, not merely equal amounts

Contemporary philosopher Elizabeth Anderson argues that equality is not exhausted by distributing identical amounts of goods. Its purpose is to create relations in which no one is treated as a subordinate class.

From the standpoint of this relational equality, neither unconditional celebration of a female umpire nor exclusion of all women after a failure is adequate. The former makes her a symbol to be protected; the latter treats women as an inferior group.

An equal relationship is one in which people may ask one another for reasons, criticize, answer, correct themselves, and still remain members of the same shared practice.

The history of rights unfolded both outside and inside sport

No women competed in the first modern Olympic Games in 1896. At Paris in 1900, twenty-two women participated for the first time—about 2.2 percent of 997 athletes. More than a century later, Paris 2024 achieved equal athlete quota places for men and women.

In Japan, women's suffrage was recognized in 1945, and Article 14 of the Constitution, which took effect in 1947, prohibited discrimination based on sex. Japan ratified the Convention on the Elimination of All Forms of Discrimination against Women in 1985, and enacted the Basic Act for a Gender-Equal Society in 1999.

But equality in law is not the same thing as gaining experience in every field.

According to the Japan High School Baseball Federation, roughly twenty female umpires were registered with prefectural federations in 2024. That year, seven women attended the national umpiring clinic for the first time. In the summer of 2026, five women were appointed to the national Koshien tournament's umpiring staff for the first time.

These facts do not fit the story that inexperienced women were suddenly placed at Koshien for the sake of equality alone. Sato had umpired high-school baseball since 2018 and had experience at an under-18 international event and at the national girls' high-school baseball championship. Her appointment followed national training and practical instruction.

Experience does not guarantee freedom from error. The point is that the success or failure of inclusion should be judged through continuing data and the same standards applied to all umpires, not through one symbolic moment.

6. Are men and women good at different things?

We now reach the question that cannot be avoided:

Are there real differences in ability between men and women?

The answer is: average differences exist in some domains, but that fact does not immediately determine any individual's fitness.

Some domains of physical performance show clear average differences

Among adults, post-pubertal differences in hormones, muscle mass, skeletal structure, and cardiorespiratory function mean that male populations, on average, outperform female populations in sports that heavily depend on strength, power, speed, and aerobic endurance. Research reviews place the performance gap at roughly 10 to 30 percent depending on the event.

There is no need to deny physical sex differences. One reason for women's sporting categories is to create fair physical conditions of competition.

But the ability to throw a fast pitch as a player is not the same as the ability to observe it and make a rules-based decision.

Umpires do need the speed to reach the proper position, the endurance to concentrate for long periods, and tolerance of heat. Physical condition is therefore not irrelevant. What the job requires, however, is not “being male,” but meeting the necessary standards for running, endurance, reaction, and positioning. Those qualities can be measured more directly than sex.

Psychological and cognitive distributions largely overlap

Psychologist Janet Hyde reviewed forty-six meta-analyses and reported that 78 percent of effect sizes for psychological gender differences were small or close to zero. This does not mean that no psychological sex differences exist. Average differences have been reported in some spatial tasks, verbal tasks, aggression, and other domains, and their size varies with age and context.

The important point is that the male and female distributions overlap extensively on many cognitive traits.

Even if men score slightly higher on average on a particular spatial test, it does not follow that every man scores above every woman. Many high-scoring women outperform many men, and many high-scoring men outperform many women.

If we can directly examine a candidate's practical test, call accuracy, rules examination, and experience, there is less reason to rely on gender as a rough proxy.

“Women are empathetic, so they make good umpires” is also risky

In answering sexism, it may be tempting to say that women are calmer, notice finer details, or are more empathetic and therefore better suited to umpiring.

But this too can become an essentialism that confines women to a single nature. Positive stereotypes, no less than negative ones, can exclude individuals who do not fit them.

Women are entitled to qualify as umpires not because they possess some uniquely feminine virtue, but because individuals with the required abilities should be able to compete and serve regardless of gender.

What does science say about officiating ability?

Officiating is not one ability.

Component What can be said about sex differences A fair way to evaluate it
Acceleration, running, endurance Some physical abilities show a higher male average Directly test the requirements of the job
Gaze and positioning Strongly shaped by experience, sport knowledge, and training Evaluate with video and tracking
Rules knowledge Learned professional knowledge Use the same written and case-based tests
Decision accuracy No evidence that sex alone determines it Compare many calls under similar conditions
Game management and explanation Shaped by experience and communication Assess observable behavior
How authority is received Bias can cause female authority to be doubted Separate audience prejudice from the official's ability

A 2025 study of handball referees found no significant sex difference in gaze behavior or decision accuracy. Another study comparing male and female referees in men's soccer found no significant differences in major match indicators such as shots, passes, fouls, and offsides.

These studies concern limited sports and samples. They do not prove that every sport is entirely free of sex differences, and large-scale comparisons of male and female baseball umpires remain scarce.

An honest conclusion therefore lies between two extremes:

  • Science has not established that female and male umpires must be perfectly identical in every respect.
  • One bad call by one woman has not established that women are unfit to umpire.

More instructively, an analysis of more than three million Major League Baseball calls indicates that umpires' perceptual skill improves through monitoring, training, and feedback. In thinking about officiating ability, how someone accumulates experience and learns from mistakes is a more direct variable than sex.

7. Aristotle and the question of a “relevant difference”

Aristotle himself defended a male-dominated social order that must be criticized from a contemporary standpoint. Yet his theory of justice leaves us a useful question.

If equals should be treated equally and relevantly different cases differently, we must ask: Which differences are relevant to the purpose at hand?

In a 100-meter race, post-pubertal physical differences directly affect results. A sex category can therefore have a rational role.

What about judging whether a catch was completed at second base?

The relevant qualities are line of sight, angle, distance, rules knowledge, reaction, and coordination with the other umpires. Unless sex can be shown to represent those qualities with sufficient accuracy, it has little relevance as a standard for officiating ability.

Fairness does not ignore every difference. It uses differences relevant to the purpose and refuses to turn irrelevant differences into rank.

8. Do not turn a group average into a verdict on an individual

Debates about sex differences become confused above all when descriptions of groups are mixed with judgments about individuals.

Suppose the average score of one thousand men on some ability is higher than the average score of one thousand women. That fact may provide probabilistic information about two randomly selected people. But if the two individuals before us have already taken a practical test and we know their results, we should examine those results rather than their sex.

Sex is a rough proxy used when no individual information is available. Candidates for elite umpiring, moreover, are not random members of the population. They are a selected group with years of experience, training, and recommendation. Population averages cannot simply be carried over to selected individuals.

Gender equality can therefore be restated as a statistical ethic:

When an ability can be measured directly, do not use sex as its proxy.

That is not a rejection of science. It is a scientific preference for the more precise evidence in front of us over a coarse category.

9. Does actively including women violate meritocracy?

The appointment of female umpires also prompts this objection:

Did the desire to demonstrate equality cause the organizers to choose women at the expense of competence?

This question should not simply be dismissed as sexist. Because an umpire's decisions can alter players' fates, asking whether selection standards are sound is legitimate.

But testing the concern requires at least the following information:

  • What qualifications and nomination conditions are shared by male and female candidates?
  • How are practical skill, rules knowledge, fitness, and game management evaluated?
  • How much experience does each candidate have, and at what level?
  • What do accuracy data and training results show across all umpires?
  • How will appointments be reviewed, and what retraining follows?

The fact that a woman was selected is not itself evidence that standards were lowered. Public information shows that the women appointed in 2026 had substantial umpiring histories or international experience and received practical training at the national clinic.

At the same time, the historic character of an appointment does not prove competence. The federation should not hide standards in an attempt to protect women. It should make the common system of selection, evaluation, and retraining as clear as possible.

Active inclusion can coexist with meritocracy under these conditions:

  1. Widen the entrance to the candidate pool.
  2. Measure abilities relevant to the work by the same standards.
  3. Use training to address historically unequal access to experience.
  4. Continue evaluating all appointees by the same standards.
  5. Do not exclude an entire group after one person's failure.

Meritocracy is not the preservation of a closed network or inherited custom. It is the attempt to find ability in a broader population by using more direct measures.

10. From infallible authority to corrigible authority

An umpire has authority. An umpire stops play, decides safe and out, and may change the result of a game.

But human perception inevitably errs. Fast plays, blind spots, crowd pressure, expectation, fatigue, and the previous call all influence judgment. Male and female umpires alike remain inside these human conditions.

If good authority is defined as a person who never errs, authority acquires an incentive to hide error. Changing an announced decision looks like weakness.

If good authority is instead defined as the ability to discover error, explain it, and correct it appropriately, consultation and video review do not damage authority. They help sustain its credibility.

The sixth-inning decision was wrong. The umpire herself asked for a conference, and it was corrected. Both facts can be assessed at once.

It was a missed call, so it requires review. She sought its correction herself, so that professional response also matters.

We do not have to erase either half.

11. Six tests for fair criticism on social media

When criticizing a sports decision—not only one made by a female umpire—we can ask six questions.

1. The subject test

Is the grammatical and moral subject “this call,” or “women”?

2. The evidence test

Is the criticism based on video, rules, positioning, or an official explanation?

3. The comparison test

Would the critic respond with the same force if a male umpire made the same mistake?

4. The generalization test

Does the argument jump from one act by one person to the essence of a whole group?

5. The proportionality test

Is the demanded response proportionate to the error? Does it call for excluding an entire group when video review or retraining would address the problem?

6. The recovery test

Does it leave open the possibility that someone who admits and learns from an error may officiate again?

Criticism that passes these six tests can be severe without being discriminatory. Conversely, even polite language is not fair technical criticism if it begins with the conclusion that women are unsuited to the role.

12. Why does gender equality matter?

The significance of gender equality is not exhausted by increasing the number of women.

1. It makes each person the author of their life

In a society that assigns work and ways of living by sex in advance, people cannot author their own lives. Equality returns not only the freedom to try and succeed, but also the freedom to try and fail.

2. It prevents society from overlooking talent

If half the population is excluded from a candidate pool by a rough stereotype, society loses capable people. Equality is not charity; it can also improve the accuracy of selection.

3. It moves authority from status to ability

Authority should rest not on an image of the “right body for an umpire,” the “right voice for a leader,” or masculine decisiveness, but on actual accuracy, rules knowledge, explanation, and coordination. That also makes the evaluation of male umpires more transparent.

4. It releases men from fixed roles too

Gender equality is not only a benefit to women. It also frees men from the demand always to be strong and decisive, to prioritize work over family, and to occupy authority. Loosening roles assigned by sex expands everyone's options.

5. It builds institutions that turn failure into learning

An organization that treats a minority member's mistake as a reason for exclusion encourages concealment. An organization that records everyone's decisions, reviews them on video, corrects them, and retrains officials becomes stronger across gender.

13. What this incident should really make us ask

If the incident ends with “female umpires are no good,” nothing has been learned.

If it ends instead with “this was historic, so no criticism is allowed,” responsibility to the players has also been lost.

The useful questions are more concrete:

  • Why was the incomplete catch missed?
  • Where was the best position from which to observe it?
  • At what point could another umpire have assisted?
  • Was the video review explained clearly enough to spectators?
  • How often do comparable errors occur among all umpires, male and female?
  • Does the rate fall after training?

These questions can help the players, the umpires, and the next game.

The first women appointed to umpire at Koshien need not be made saints. Nor must they be made symbols of failure. They can be treated as individual professionals: examine the call, recognize a good corrective response, demand improvement, and evaluate the next assignment by the same standards.

Gender equality is not a doctrine claiming that male and female bodies, minds, and statistical distributions are perfectly identical.

Equality does not erase averages. It refuses to turn an average into a verdict on a person.

One missed call proves that human judgment is fallible. To prove anything more requires evidence adequate to the additional claim.

And if fallible human beings are to create a fair game together, they need institutions that can examine, explain, and correct decisions—not declarations based on sex.


NOW IN QUESTION

When we see one mistake by a female umpire, would we truly direct the same words at a male umpire who made the same mistake?

And when we say that we support her because she is a woman, or refuse to criticize her because she is a woman, are we respecting her as an individual professional?

Does the equality we want require everyone to be made to look the same? Or does it ask us to evaluate each person through their own conduct and ability, even when differences exist?

FAQ

Q1. Was there actually a missed call in this game?

The sixth-inning out call was changed to safe after the umpiring crew determined that the catch had not been completed, and Sato herself described it as her mistake. The seventh-inning steal remained an out after video review. It is therefore inaccurate to state with equal confidence that there were “two bad calls.”

Q2. Is criticizing a female umpire's missed call sexist?

No. A particular call should be criticized when video and the rules support that criticism. It becomes sexist when one call is generalized into the claim that women are unfit to umpire, or when a different standard is applied to male officials.

Q3. Are there no ability differences between men and women?

It depends on the domain. Strength, speed, and power show clear average differences. Average differences in many psychological and cognitive traits are small and the distributions overlap greatly, while particular tasks do show differences. Group averages must be kept separate from the ability of the individual before us.

Q4. If men have greater average physical ability, are men also better suited to umpire boys' baseball?

The running and endurance required of an umpire can be measured through practical standards rather than sex. An umpire need not match the pitching or hitting ability of the male players. The rational approach is to establish common job-related fitness standards and assess each individual.

Q5. Does actively appointing women amount to reverse discrimination or disregard for merit?

That depends on the method. Widening the candidate pool and training opportunities while evaluating job-relevant abilities by common standards is compatible with merit. Whether standards were lowered should be judged from the standards and results, not from the mere fact that a woman was chosen. The institution owes the public transparency.

Q6. Is praising the correction of the call a softer standard for a woman?

The original error and the appropriateness of the correction can both be true. A male umpire who recognizes an error and asks for a conference can also be credited for that response. The key is to evaluate the conduct in the same way regardless of gender.

Q7. What, exactly, should gender equality make equal?

It does not make everyone's ability or results identical. It makes opportunities to try, relevant evaluation standards, accountability, learning, and second chances equal—and rejects treating someone as a subordinate class because of sex.

Fact-checking caveats

  • This essay was prompted by the linked X post, but it does not claim to represent the intent of every poster or every reaction on social media.
  • The account of the game was checked against public information from the Japan High School Baseball Federation, The Mainichi, High School Baseball Dot Com, and Sports Hochi, among others.
  • Sato acknowledged the sixth-inning mistake and the call was corrected after consultation. The seventh-inning decision stood after video review.
  • The philosophers discussed here did not write about this game or Sato. Their ideas are being applied to a contemporary case.
  • Research on sex differences concerns group averages. It is not a basis for deciding an individual's ability or fitness from sex alone.
  • “Gendering a bad call” and “representative tax” are analytical concepts proposed in this essay.

References

The game and high-school baseball umpiring

Research on sex differences and officiating

The history and institutions of gender equality

Philosophy

The philosophical discussions above summarize primary texts and English-language reference materials rather than reproducing them at length.

SHARE

Share this article

Carry the question into another conversation.

DISCUSSION

Discuss this question

Write in Japanese or English. Comments are translated automatically and shared across both versions of the article.

0/2000 characters

Loading comments…