Judging评审机制
Four criteria, twenty points
四项标准,总分 20
Every application is scored on four equally weighted criteria, each 1.0–5.0, summing to 20. No single strength can carry a weak criterion — the composite is a sum, not a max.
每份申请按四项等权标准评分,每项 1.0–5.0 分,总分 20。任何单项优势都无法弥补短板——总分是求和,不是取最大值。
4 × 5.0 = 204 项 × 5.0 = 20 分
Entry Form 申请材料
25% · 5.0Academics, transcript rigor, recommendations, activities and leadership.学业成绩、课程难度、推荐信、课外活动与领导力。
Scientific Merit 科学价值
25% · 5.0Research validity, methodology sophistication, experimental design, quality of analysis.研究的有效性、方法的深度、实验设计与分析质量。
Student Contribution 学生贡献
25% · 5.0Independence, initiative, originality — how much of the core work is demonstrably the student's own.独立性、主动性、原创性——核心工作有多少可以证明是学生本人完成的。
Scientific Potential 综合科学潜力
25% · 5.0Scientific ability and future promise, read largely through the essays and communication quality.科学能力与未来潜力,主要通过短文与表达质量来判断。
The Mechanics筛选机制
The "On The Table" cut is a union of two lenses「On The Table」是两个镜头的并集
From our reconstruction of a full 463-project scored docket (2,471 entries that cycle): a project reaches the scored table if it lands in the Top 400 by Z-score OR the Top 350 by raw score. The two rankings correlate at only r ≈ 0.18, so which lens a project is measured through matters enormously.
根据我们对一份完整的 463 项已评分名单的重建(该赛季共 2,471 份申请):项目进入评分名单的条件是 Z 分排名前 400,或原始分排名前 350。两种排名的相关系数仅 r ≈ 0.18——用哪个镜头衡量,结果差别巨大。
400
Top 400 by Z-scoreZ 分排名前 400
best relative to your category's evaluator pool相对于本类别评审池的最强者
350
Top 350 by raw score原始分排名前 350
best on the absolute 20-point scale20 分制绝对分数的最强者
463
"On The Table"进入评审桌
~19% of 2,471 entrants → 300 Scholars → 40 Finalists约占 2,471 人的 19% → 300 学者 → 40 决赛
The Z-score exists because raw scoring bars differ by category — Chemistry's mean raw score on the table was 17.7 while Behavioral Sciences' was 15.5. Evaluator pools grade differently, and Z neutralizes that.
Z 分之所以存在,是因为各类别的打分基准不同——名单上化学类的原始分均值为 17.7,而行为科学类只有 15.5。不同评审池宽严不一,Z 分正是为抹平这一差异。
r ≈ 0.18
Correlation between the two rankings两种排名之间的相关系数
~3 pts
The band the whole scored tier fits in (mean 16.7/20)整个评分梯队被压缩的窄带(均分 16.7/20)
17.7 vs 15.5
Mean raw bar, Chemistry vs Behavioral Sciences化学类与行为科学类的原始分均值之差
23
High-scoring rows parked pending integrity screens高分却被暂扣、等待诚信筛查的条目
- Small differences move you far. The scored tier is compressed into roughly a three-point band, so a fraction of a point moves a project dozens of ranks.微小的分差就能挪动很远。整个评分梯队被压缩在约 3 分的窄带内,零点几分的差距就能移动几十个名次。
- The category lens matters. You are scored relative to your category's evaluator pool — a 16.5 in one category can be more selective than a 16.5 in another.类别镜头很重要。你的得分是相对于本类别评审池而言的——某类别的 16.5 分可能比另一类别的 16.5 分更「值钱」。
- A high score does not guarantee survival. Integrity screens — plagiarism detection, AI checks, undisclosed conflicts, team-work violations — run after scoring, and high-scoring entries do get parked or removed there.高分并不保证晋级。查重、AI 检测、未披露的利益冲突与团队合作违规等诚信筛查在评分之后进行,确实有高分申请在这一关被暂扣或除名。
- Scoring carries real variance. The two-lens union is the Society's own hedge against evaluator severity and category difficulty — treat your score as measured through a noisy instrument, and maximize what you control.评分带有真实方差。双镜头并集正是主办方为对冲评审宽严与类别难度而设的机制——把分数当作一台有噪声的仪器的读数,把力气花在你能控制的事上。
Next: what the full docket reveals about who reaches the table — The Data →
下一步:完整名单揭示了谁能上桌——数据洞察 →