<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Data Analysis | Sumin Han</title><link>https://smhanlab.com/tags/data-analysis/</link><atom:link href="https://smhanlab.com/tags/data-analysis/index.xml" rel="self" type="application/rss+xml"/><description>Data Analysis</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Wed, 07 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>https://smhanlab.com/media/icon_hu3d03a01dcc18bc5be0e67db3d8d209a6_9723_512x512_fill_lanczos_center_3.png</url><title>Data Analysis</title><link>https://smhanlab.com/tags/data-analysis/</link></image><item><title>Eight Hidden Reasons Seoul Moves: Latent Travel Motives from Hourly OD Data</title><link>https://smhanlab.com/post/20261007-seoul-mobility-motives/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate><guid>https://smhanlab.com/post/20261007-seoul-mobility-motives/</guid><description>&lt;p>&lt;strong>TL;DR.&lt;/strong> I took a year of hourly trips between 481 districts of the Seoul metropolitan area (Seoul Metropolitan Government and KT mobility data, 2024) and asked whether a few hidden travel motives could explain them. The model only ever saw flow counts. Everything else in the data, including trip purpose and nationality, was kept as an answer key.&lt;/p>
&lt;ol>
&lt;li>&lt;strong>Six motives (about 6,000 parameters)&lt;/strong> explain held-out months as well as a static OD model with &lt;strong>463,000 parameters&lt;/strong>. With &lt;strong>eight motives the error is 12% lower&lt;/strong>.&lt;/li>
&lt;li>Motives need their own daily rhythms. Forcing them to share one makes the error &lt;strong>47% worse&lt;/strong>.&lt;/li>
&lt;li>The daily strength of the commuting motives moves with &lt;strong>subway smart-card counts&lt;/strong> at a correlation of &lt;strong>0.66–0.77&lt;/strong>. The long-distance motive sits at 0.16.&lt;/li>
&lt;li>On the days of the &lt;strong>Seoul International Fireworks Festival and two rallies in Yeouido&lt;/strong>, the model leaves &lt;strong>+44% to +56%&lt;/strong> unexplained arrivals in Yeouido-dong. Ordinary Saturdays show &lt;strong>−23% and −32%&lt;/strong>.&lt;/li>
&lt;li>What did not work: recovering the held-out trip-purpose label beats a time-of-day-only guess by just &lt;strong>1.6 percentage points&lt;/strong>, and no &amp;ldquo;tourist&amp;rdquo; motive appeared, because short-term foreigners move almost exactly like Korean nationals (hourly correlation &lt;strong>0.976&lt;/strong>).&lt;/li>
&lt;li>A side result: Seoul commute distances &lt;strong>did not grow&lt;/strong> in 2023–2026 (mean &lt;strong>11.11 km → 10.88 km&lt;/strong>).&lt;/li>
&lt;/ol>
&lt;p>The full write-up is a working paper: &lt;strong>&lt;a href="han2026seoulmotives_en.pdf">PDF (English)&lt;/a>&lt;/strong>. The original Korean draft is also available: &lt;strong>&lt;a href="han2026seoulmotives_ko.pdf">Korean version (PDF)&lt;/a>&lt;/strong>. If you prefer to watch, here is a five-minute video version:&lt;/p>
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/YmkSxE1oW0w?autoplay=0&amp;controls=1&amp;end=0&amp;loop=0&amp;mute=0&amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
>&lt;/iframe>
&lt;/div>
&lt;hr>
&lt;h2 id="why-ask-this-question">Why ask this question?&lt;/h2>
&lt;p>Mobility data has become very rich. Telecom-based datasets now tell us, every day, how many people moved from one district to another and at what hour. What they do not tell us directly is &lt;em>why&lt;/em>. Travel surveys ask about purpose, but they run every few years on small samples. The purpose field in telecom data is not an answer either: it is the provider&amp;rsquo;s own rule-based guess.&lt;/p>
&lt;p>So I turned the question around. Suppose that behind all these flows there are a few hidden motives, and each motive has a simple character: where its trips tend to start, where they tend to end, at what hour they happen, and how far they go. Can we find those motives from the flows alone? And if we can, do they agree with information the model has never seen?&lt;/p>
&lt;p>Before training anything, I wrote down six hypotheses, each with an expected result and a reason. I report the ones that failed alongside the ones that held.&lt;/p>
&lt;h2 id="the-data">The data&lt;/h2>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Source&lt;/th>
&lt;th>What it gives&lt;/th>
&lt;th>How I used it&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Seoul Metropolitan Mobility (Seoul Metropolitan Government &amp;amp; KT, OA-22300)&lt;/td>
&lt;td>hourly trips between districts, with purpose, Korean/foreign status, distance&lt;/td>
&lt;td>main data: 2024 for the model, 2023–2026 for checks&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Subway smart-card counts&lt;/td>
&lt;td>daily entries and exits, Seoul-wide&lt;/td>
&lt;td>independent test of the motives&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>TOPIS road detectors&lt;/td>
&lt;td>daily road volume and mean speed&lt;/td>
&lt;td>independent test&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Known event days&lt;/td>
&lt;td>fireworks festival, rallies, holidays, heavy snow&lt;/td>
&lt;td>test of what the model cannot explain&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The network has &lt;strong>481 nodes&lt;/strong>: 426 Seoul districts (&lt;em>dong&lt;/em>), 54 cities and counties in Gyeonggi and Incheon, and one node for the rest of the country. I kept every flow that starts or ends in Seoul (43.4% of all 2024 flows). The model was trained on January–September 2024 and tested on October–December. In the training months, Seoul-related flows are about 26.3 million trips on a weekday and 21.1 million on a weekend day.&lt;/p>
&lt;h2 id="the-model-in-words">The model, in words&lt;/h2>
&lt;p>Each motive is a small recipe with four ingredients:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>where from&lt;/strong>: a distribution over origin districts&lt;/li>
&lt;li>&lt;strong>where to&lt;/strong>: a distribution over destination districts&lt;/li>
&lt;li>&lt;strong>what time&lt;/strong>: its own 24-hour rhythm, separately for weekdays and weekends&lt;/li>
&lt;li>&lt;strong>how far&lt;/strong>: a distance scale, so that trips become rarer as the distance grows&lt;/li>
&lt;/ul>
&lt;p>The expected number of trips between two districts at a given hour is the sum of what each motive contributes. Three old ideas sit underneath. Flows are counts of people, so the model uses a Poisson likelihood. The distance part is a gravity model, the workhorse of transport research. And splitting the total into a sum of parts is the same idea as a topic model, where a document is a mix of topics.&lt;/p>
&lt;p>The model only sees counts. Trip purpose, nationality and distance labels are never used for training. Each fit takes a few seconds on one GPU.&lt;/p>
&lt;h2 id="result-1-a-few-motives-are-enough">Result 1: a few motives are enough&lt;/h2>
&lt;p>&lt;img src="fig1_ksweep.png" alt="Validation deviance versus number of motives">&lt;/p>
&lt;p>The red dashed line is a strong but blunt baseline: it memorizes the average size of every origin–destination pair (462,770 numbers) and multiplies it by one common daily curve. Six motives reach the same validation error (0.298 vs. 0.299) with 76 times fewer parameters. Eight motives are 12% better (0.263). The gray line is the same model with one shared daily rhythm, and it is much worse: 47% at eight motives.&lt;/p>
&lt;p>There is no sharp &amp;ldquo;elbow&amp;rdquo; in the curve, so the error alone cannot tell me the right number of motives. I used eight as the overview and sixteen for detail. Honest caveat: the motives are not perfectly reproducible across random seeds. Large motives come back reliably, but small ones merge and split. Read the individual motives as examples; the aggregate tests below are more trustworthy.&lt;/p>
&lt;h2 id="result-2-what-the-motives-look-like">Result 2: what the motives look like&lt;/h2>
&lt;p>&lt;img src="fig2_rhythms.png" alt="Daily rhythm of each motive">&lt;/p>
&lt;p>The eight motives, in order of size:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Motive (my interpretation)&lt;/th>
&lt;th style="text-align:right">Share of flow&lt;/th>
&lt;th>Weekday peak&lt;/th>
&lt;th>Where (from actual flows)&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Evening return, center → Gyeonggi&lt;/td>
&lt;td style="text-align:right">18.2%&lt;/td>
&lt;td>18 h&lt;/td>
&lt;td>Yeoksam 1-dong, Yeouido-dong, Jongno → Namyangju, Goyang, Bundang&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Outer-area local life&lt;/td>
&lt;td style="text-align:right">17.0%&lt;/td>
&lt;td>8 h (weekend 13 h)&lt;/td>
&lt;td>within Hanam, Goyang, Jingwan-dong&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Morning commute, Gyeonggi → center&lt;/td>
&lt;td style="text-align:right">12.6%&lt;/td>
&lt;td>7 h&lt;/td>
&lt;td>Namyangju, Bundang, Goyang → Yeoksam 1-dong, Yeouido-dong, Jongno&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Long trips and gateways&lt;/td>
&lt;td style="text-align:right">12.4%&lt;/td>
&lt;td>9 h (weekend 14 h)&lt;/td>
&lt;td>rest of country, Gonghang-dong ↔ Incheon Airport, Banpo 4-dong&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Downtown daytime&lt;/td>
&lt;td style="text-align:right">11.4%&lt;/td>
&lt;td>12 h&lt;/td>
&lt;td>within Jongno, Yeouido-dong, Hangangno-dong&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Outer Seoul → center commute&lt;/td>
&lt;td style="text-align:right">11.2%&lt;/td>
&lt;td>7 h&lt;/td>
&lt;td>Segok-dong, Jingwan-dong, Doksan 1-dong → Yeouido-dong, Yeoksam 1-dong&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Short-range evening leisure&lt;/td>
&lt;td style="text-align:right">8.8%&lt;/td>
&lt;td>12 h (weekend 19 h)&lt;/td>
&lt;td>within Yeoksam 1-dong, Hwayang-dong, Sageun-dong&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Neighborhood nights&lt;/td>
&lt;td style="text-align:right">8.3%&lt;/td>
&lt;td>8 h (weekend 21 h)&lt;/td>
&lt;td>within Sinchon-dong, Seogyo-dong, Heukseok-dong&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The morning commute and evening return connect the same business districts with the same Gyeonggi cities, in opposite directions and at opposite hours. The long-trips motive, built around airports and the Express Bus Terminal, is 1.32 times stronger on weekends than on weekdays. The smaller motives look like local clusters around one or two neighborhoods.&lt;/p>
&lt;p>&lt;img src="day_flower.png" alt="The eight motives around the 24-hour clock">&lt;/p>
&lt;p>Laid out on a map, the same roads switch motive with the hour: at 07:00 gold commute lines converge on the center, and at 18:00 red return lines fan back out.&lt;/p>
&lt;p>&lt;img src="fig3_maps.png" alt="Modeled flows at three times of day">&lt;/p>
&lt;p>(The node positions on the maps are approximate, because I did not have district boundary polygons. They are used only for drawing.)&lt;/p>
&lt;h2 id="result-3-the-motives-move-with-an-independent-sensor">Result 3: the motives move with an independent sensor&lt;/h2>
&lt;p>The strongest test uses data that has nothing to do with telecom records. I froze the motives and estimated only how strong each one was on each day of 2024. Then I compared those daily strengths with subway smart-card counts, after removing the usual weekday pattern from both.&lt;/p>
&lt;p>&lt;img src="fig4_sensors.png" alt="Correlation between daily motive strength and independent sensors">&lt;/p>
&lt;p>The commuting motives move closely with the subway: &lt;strong>0.77, 0.73 and 0.66&lt;/strong>. The long-trips motive barely does (&lt;strong>0.16&lt;/strong>), which makes sense for airport and intercity travel. Road volume shows the same ordering but weaker (at most 0.53). Because the two measurements are independent, rising and falling together on the same days is good evidence that the motives reflect real behavior and not just a convenient factorization.&lt;/p>
&lt;h2 id="result-4-what-the-model-cannot-explain-on-special-days">Result 4: what the model cannot explain on special days&lt;/h2>
&lt;p>If the motives capture the ordinary city, unusual gatherings should show up in what is left over.&lt;/p>
&lt;p>&lt;img src="fig5_events.png" alt="Excess arrivals in Yeouido-dong on event days, and motive strength on holidays">&lt;/p>
&lt;p>On the day of the Seoul International Fireworks Festival (October 5), Yeouido-dong received &lt;strong>237,074 more arrivals than expected (+44%)&lt;/strong>, and riverside districts nearby also lit up (Ichon 2-dong +86%). On the two rally days in Yeouido (December 7 and 14), the excess was &lt;strong>+49% and +56%&lt;/strong>. Two ordinary Saturdays, used as controls, show &lt;strong>−23% and −32%&lt;/strong>. The events barely move the daily motive strengths; they appear only in the destination residuals, which is what the design predicts.&lt;/p>
&lt;p>Holidays tell a mixed story. On Lunar New Year and Chuseok, the evening-return motive falls to about a third of a normal weekday and long trips rise by 1.6–1.9 times. But one commuting motive hardly falls at all (0.99 and 0.95), so my prediction that commuting would drop below 30% was wrong.&lt;/p>
&lt;p>One finding I did not expect: the model picks out &lt;strong>de facto days off&lt;/strong> without any calendar. Commuting collapsed on Labor Day (May 1), the day after Memorial Day (June 7), September 30 before a temporary holiday, and the last days of December.&lt;/p>
&lt;h2 id="a-side-question-are-commutes-getting-longer">A side question: are commutes getting longer?&lt;/h2>
&lt;p>Housing prices in Seoul have risen, and a common assumption is that people now commute farther. The same data can check this for work trips on Tuesdays to Thursdays.&lt;/p>
&lt;p>&lt;img src="fig6_commute.png" alt="Monthly trends of commute distance">&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>January–August average&lt;/th>
&lt;th style="text-align:right">2023&lt;/th>
&lt;th style="text-align:right">2024&lt;/th>
&lt;th style="text-align:right">2025&lt;/th>
&lt;th style="text-align:right">2026&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Mean commute distance (km)&lt;/td>
&lt;td style="text-align:right">11.11&lt;/td>
&lt;td style="text-align:right">11.07&lt;/td>
&lt;td style="text-align:right">10.93&lt;/td>
&lt;td style="text-align:right">10.88&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Share ≥ 20 km&lt;/td>
&lt;td style="text-align:right">16.1%&lt;/td>
&lt;td style="text-align:right">16.1%&lt;/td>
&lt;td style="text-align:right">15.8%&lt;/td>
&lt;td style="text-align:right">15.8%&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Seoul-bound commutes from Gyeonggi/Incheon&lt;/td>
&lt;td style="text-align:right">26.2%&lt;/td>
&lt;td style="text-align:right">26.2%&lt;/td>
&lt;td style="text-align:right">25.8%&lt;/td>
&lt;td style="text-align:right">25.7%&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>There is no sign of an increase. Mean distance is drifting down by about 0.8% per year. This does not mean housing prices have no effect: three and a half years is short, moving house changes commutes slowly, and this data contains no prices at all.&lt;/p>
&lt;h2 id="what-did-not-work">What did not work&lt;/h2>
&lt;p>&lt;strong>Recovering trip purpose.&lt;/strong> I expected the motives to recover the provider&amp;rsquo;s purpose labels well, around 65–70% accuracy. They did not.&lt;/p>
&lt;p>&lt;img src="fig7_h1.png" alt="Accuracy of predicting the held-out trip purpose">&lt;/p>
&lt;p>Used as designed, one purpose mix per motive, the motives did worse than a guess based on time of day alone (57.6% vs. 60.5%). Letting each motive have a purpose mix per hour reaches 62.1%, which is only &lt;strong>1.6 percentage points&lt;/strong> above time only and far from the supervised ceiling of 64.8%. Purpose labels are rule-based guesses by the provider, mostly determined by home and work locations and time of day, and my motives added little beyond that.&lt;/p>
&lt;p>&lt;strong>Finding tourists.&lt;/strong> I expected a motive with at least twice the average share of short-term foreigners, heading to Myeong-dong or Hongdae with no morning peak. None appeared. When I looked into why, the label itself turned out to be the issue: trips labeled as short-term foreigners have almost the same hourly pattern as those of Korean nationals (correlation 0.976 on weekdays, similar rush-hour shares), and they make up 16.5% of all trips, about 4.3 million a day. That is not what tourists look like. I could not confirm what the label really captures.&lt;/p>
&lt;p>Two more misses: heavy snow in late November did not visibly delay commuting, and the night of the martial-law declaration on December 3 left no clear trace in the daily flows.&lt;/p>
&lt;h2 id="limitations">Limitations&lt;/h2>
&lt;ul>
&lt;li>Individual motives depend on the random seed (matched cosine 0.74–0.83). The large motives are stable; the small ones are less so.&lt;/li>
&lt;li>Purpose and nationality labels are the provider&amp;rsquo;s estimates, so recovering them would not mean recovering real reasons.&lt;/li>
&lt;li>The model describes the average day. Events show up only as residuals.&lt;/li>
&lt;li>Districts outside Seoul are coarse (city or county level), and trips that never touch Seoul are excluded.&lt;/li>
&lt;li>The commute-distance trend covers only 3.6 years and cannot test a link with housing prices.&lt;/li>
&lt;/ul>
&lt;h2 id="a-note-on-how-this-was-made">A note on how this was made&lt;/h2>
&lt;p>I used an AI assistant (Claude) to write the analysis and figure code, draw the figures, produce the narrated video, and write the English versions of the paper and this post from my Korean draft. Every number was computed directly from the data by that code. I chose the questions and the hypotheses, and I am responsible for the interpretation.&lt;/p>
&lt;p>If you work with telecom mobility data, or you know what the short-term-foreigner label in this dataset actually measures, I would be glad to hear from you.&lt;/p>
&lt;p>&lt;strong>Download:&lt;/strong> &lt;a href="han2026seoulmotives_en.pdf">Working paper, English (PDF)&lt;/a> · &lt;a href="han2026seoulmotives_ko.pdf">Korean version (PDF)&lt;/a> · &lt;a href="https://data.seoul.go.kr/dataList/OA-22300/F/1/datasetView.do">Seoul Metropolitan Mobility data (OA-22300)&lt;/a>&lt;/p></description></item><item><title>What 320,000 Public Reviews Say About AI Conferences</title><link>https://smhanlab.com/post/20261007-ai-conference-reviews/</link><pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate><guid>https://smhanlab.com/post/20261007-ai-conference-reviews/</guid><description>&lt;p>&lt;strong>TL;DR.&lt;/strong> I downloaded every public review on OpenReview for the three largest machine-learning conferences: &lt;strong>84,212 papers and 322,903 reviews&lt;/strong> from ICLR 2018–2026, NeurIPS 2021–2025, and ICML 2025–2026. Six things stood out.&lt;/p>
&lt;ol>
&lt;li>ICLR submissions grew &lt;strong>19-fold in eight years&lt;/strong>. Almost a third of 2026 submissions were withdrawn or desk-rejected before a decision.&lt;/li>
&lt;li>One in five ICLR submissions now has &lt;strong>&amp;ldquo;LLM&amp;rdquo; or &amp;ldquo;language model&amp;rdquo; in the title&lt;/strong>, up from 1% in 2018. Those papers are accepted at &lt;strong>exactly the average rate&lt;/strong>.&lt;/li>
&lt;li>The most common weakness reviewers write down is &lt;strong>missing or weak baselines (34%)&lt;/strong>, ahead of novelty (22%).&lt;/li>
&lt;li>At ICLR 2026, a &lt;strong>mean score of 5.0 meant a 54% chance&lt;/strong> of acceptance. Each half point near the threshold was worth about &lt;strong>25 percentage points&lt;/strong>.&lt;/li>
&lt;li>Only &lt;strong>5 of 1,209&lt;/strong> decided papers without any author reply were accepted.&lt;/li>
&lt;li>The word &lt;strong>&amp;ldquo;delve&amp;rdquo;&lt;/strong> appeared in reviews nine times more often in 2024 than in 2022, and then dropped back.&lt;/li>
&lt;/ol>
&lt;p>The full write-up with methods and limitations is available as a working paper: &lt;strong>&lt;a href="han2026aireviews.pdf">PDF&lt;/a>&lt;/strong>. If you prefer to watch, here is a five-minute video version:&lt;/p>
&lt;div style="position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;">
&lt;iframe allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen="allowfullscreen" loading="eager" referrerpolicy="strict-origin-when-cross-origin" src="https://www.youtube.com/embed/Ej-7fvowj1U?autoplay=0&amp;controls=1&amp;end=0&amp;loop=0&amp;mute=0&amp;start=0" style="position: absolute; top: 0; left: 0; width: 100%; height: 100%; border:0;" title="YouTube video"
>&lt;/iframe>
&lt;/div>
&lt;hr>
&lt;h2 id="why-look-at-reviews-at-all">Why look at reviews at all?&lt;/h2>
&lt;p>Every researcher in machine learning has a peer-review story. Usually it involves a reviewer who seemed to have read a different paper. The conversation tends to stay anecdotal, but it no longer has to. &lt;a href="https://openreview.net">OpenReview&lt;/a> publishes reviews, scores, author responses, reviewer follow-ups, and final decisions for many major venues. At ICLR, this includes rejected and withdrawn papers.&lt;/p>
&lt;p>So I asked four questions that I, and many graduate students I know, actually care about:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>What are people submitting&lt;/strong>, and how fast is that changing?&lt;/li>
&lt;li>&lt;strong>What do reviewers object to&lt;/strong> most often?&lt;/li>
&lt;li>&lt;strong>What score is &amp;ldquo;enough&amp;rdquo;?&lt;/strong>&lt;/li>
&lt;li>&lt;strong>Does the rebuttal matter?&lt;/strong>&lt;/li>
&lt;/ul>
&lt;p>Along the way I also looked at the reviews themselves: how long they are, and whether they show the vocabulary fingerprints of LLM-written text.&lt;/p>
&lt;h2 id="the-data">The data&lt;/h2>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Venue&lt;/th>
&lt;th>Years&lt;/th>
&lt;th style="text-align:right">Papers&lt;/th>
&lt;th style="text-align:right">Reviews&lt;/th>
&lt;th>What is public&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>ICLR&lt;/td>
&lt;td>2018–2026&lt;/td>
&lt;td style="text-align:right">55,472&lt;/td>
&lt;td style="text-align:right">209,348&lt;/td>
&lt;td>every submission, including rejected and withdrawn&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>NeurIPS&lt;/td>
&lt;td>2021–2025&lt;/td>
&lt;td style="text-align:right">18,763&lt;/td>
&lt;td style="text-align:right">75,254&lt;/td>
&lt;td>accepted + rejected papers whose authors opted in&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>ICML&lt;/td>
&lt;td>2025–2026&lt;/td>
&lt;td style="text-align:right">9,977&lt;/td>
&lt;td style="text-align:right">38,301&lt;/td>
&lt;td>accepted + rejected papers whose authors opted in&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;strong>Total&lt;/strong>&lt;/td>
&lt;td>16 venue-years&lt;/td>
&lt;td style="text-align:right">&lt;strong>84,212&lt;/strong>&lt;/td>
&lt;td style="text-align:right">&lt;strong>322,903&lt;/strong>&lt;/td>
&lt;td>&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>I used the OpenReview API to download each submission&amp;rsquo;s title, abstract, keywords, and every reply: reviews, meta-reviews, decisions, author comments, reviewer comments, and withdrawal notices. I did not download any PDFs. Collection finished on October 7, 2026.&lt;/p>
&lt;p>&lt;strong>Why mostly ICLR?&lt;/strong> At NeurIPS and ICML, about 95% of the papers with public reviews are accepted. That reflects the publication policy, not the real acceptance rate. Comparing accepted and rejected papers there would be misleading, so every statistic that needs both groups uses ICLR, where the whole population is public.&lt;/p>
&lt;h2 id="1-growth-and-a-third-of-submissions-that-are-never-decided">1. Growth, and a third of submissions that are never decided&lt;/h2>
&lt;p>&lt;img src="fig1_growth.png" alt="ICLR submissions per year and outcome of each submission">&lt;/p>
&lt;p>ICLR received &lt;strong>1,018 submissions in 2018 and 19,814 in 2026&lt;/strong>. Growth has accelerated rather than slowed: the last two cycles grew by 1.6× and then 1.7×.&lt;/p>
&lt;p>The acceptance share has stayed inside a narrow 26–33% band the whole time. What changed is the part of the bar that never reaches a decision. In 2026, &lt;strong>26.3% of submissions were withdrawn and 4.6% were desk-rejected&lt;/strong>. Desk rejections went from 53 in 2024 and 71 in 2025 to &lt;strong>908 in 2026&lt;/strong>. Almost one submission in three now leaves the process before a final decision.&lt;/p>
&lt;h2 id="2-topics-llms-everywhere-and-gans-nowhere">2. Topics: LLMs everywhere, and GANs nowhere&lt;/h2>
&lt;p>&lt;img src="fig2_topics.png" alt="Share of ICLR submissions whose title contains each keyword">&lt;/p>
&lt;p>I matched simple keywords against paper &lt;strong>titles&lt;/strong>. A title is a decent proxy for how authors want their work to be seen.&lt;/p>
&lt;ul>
&lt;li>&lt;strong>LLMs.&lt;/strong> Titles with &amp;ldquo;LLM&amp;rdquo; or &amp;ldquo;language model&amp;rdquo; stayed at about 1% through 2022. They reached 2.9% in 2023, &lt;strong>11.5% in 2024&lt;/strong>, and &lt;strong>21.4% in 2026&lt;/strong>, so roughly one submission in five. The 2023→2024 jump is four-fold. The ICLR 2024 deadline (September 2023) was the first after ChatGPT&amp;rsquo;s public release in November 2022.&lt;/li>
&lt;li>&lt;strong>GAN → diffusion.&lt;/strong> GAN titles fell from 5.8% in 2018 to 0.1% in 2026. Diffusion and flow matching rose from 0.8% in 2022 to 7.0% in 2025.&lt;/li>
&lt;li>&lt;strong>What is growing now.&lt;/strong> &lt;em>Reasoning&lt;/em> went from 1.7% to 7.9% between 2024 and 2026 (×4.6), and &lt;em>agent&lt;/em> from 2.0% to 6.2% (×3.1). The field&amp;rsquo;s attention seems to be moving from the models themselves to what they can reason about and do.&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Does riding the wave help?&lt;/strong> No. Among decided ICLR 2026 papers, LLM-titled papers were accepted at &lt;strong>39.2%&lt;/strong>, against &lt;strong>39.0% overall&lt;/strong>. Diffusion papers did a little better (45.8%), reasoning 44.0%, and agent papers a little worse (36.6%). A hot topic mostly means more competition.&lt;/p>
&lt;h2 id="3-what-reviewers-actually-object-to">3. What reviewers actually object to&lt;/h2>
&lt;p>&lt;img src="fig3_weaknesses.png" alt="Share of ICLR 2026 reviews whose weakness section matches each theme">&lt;/p>
&lt;p>ICLR 2026 reviews have a separate &lt;em>weaknesses&lt;/em> field. I matched twelve themes against it with keyword patterns. A single review can match several themes.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Theme&lt;/th>
&lt;th style="text-align:right">Share of reviews&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Baselines / comparisons&lt;/td>
&lt;td style="text-align:right">&lt;strong>33.5%&lt;/strong>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Scalability / compute cost&lt;/td>
&lt;td style="text-align:right">&lt;strong>31.1%&lt;/strong>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Theory / proofs&lt;/td>
&lt;td style="text-align:right">22.4%&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Clarity / writing&lt;/td>
&lt;td style="text-align:right">22.3%&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Novelty / incremental&lt;/td>
&lt;td style="text-align:right">22.1%&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Generalization&lt;/td>
&lt;td style="text-align:right">14.1%&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>Ablations&lt;/td>
&lt;td style="text-align:right">11.9%&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The usual story is that reviewers kill papers for &amp;ldquo;lack of novelty&amp;rdquo;. In the written weaknesses, though, the most common question is &lt;strong>&amp;ldquo;is it better than what already exists, and at what cost?&amp;rdquo;&lt;/strong>. It comes up about one and a half times as often as &amp;ldquo;is it new?&amp;rdquo;. Before submitting, it is worth checking that you have the strongest baselines, a fair comparison, and an honest accounting of compute.&lt;/p>
&lt;h2 id="4-score-cutoffs-the-steepest-part-of-the-curve">4. Score cutoffs: the steepest part of the curve&lt;/h2>
&lt;p>&lt;img src="fig4_cutoff.png" alt="ICLR 2026 acceptance rate by mean score, and by rebuttal outcome">&lt;/p>
&lt;p>The average score is &lt;strong>5.40 for accepted&lt;/strong> and &lt;strong>3.97 for rejected&lt;/strong> ICLR 2026 papers (out of 10). The left panel shows the acceptance rate of all 13,697 decided papers, grouped by mean score rounded to the nearest 0.5:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Mean score&lt;/th>
&lt;th style="text-align:right">4.0&lt;/th>
&lt;th style="text-align:right">4.5&lt;/th>
&lt;th style="text-align:right">&lt;strong>5.0&lt;/strong>&lt;/th>
&lt;th style="text-align:right">5.5&lt;/th>
&lt;th style="text-align:right">6.0&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>Accepted&lt;/td>
&lt;td style="text-align:right">12%&lt;/td>
&lt;td style="text-align:right">29%&lt;/td>
&lt;td style="text-align:right">&lt;strong>54%&lt;/strong>&lt;/td>
&lt;td style="text-align:right">79%&lt;/td>
&lt;td style="text-align:right">93%&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>So &lt;strong>5.0 is roughly a coin flip&lt;/strong>, and between 4.5 and 5.5 each half point moves the odds by about &lt;strong>25 percentage points&lt;/strong>.&lt;/p>
&lt;p>ICLR 2026 scores are 0, 2, 4, 6, 8, or 10. With four reviewers, &lt;strong>one reviewer changing a 4 to a 6 raises the mean by exactly 0.5&lt;/strong>. Near the threshold, convincing a single person is worth a quarter of your chance of acceptance.&lt;/p>
&lt;p>(Scores in this analysis are the latest version on OpenReview, so they already include any post-rebuttal changes. Score scales differ between years, so these cutoffs apply to ICLR 2026 only.)&lt;/p>
&lt;h2 id="5-rebuttals">5. Rebuttals&lt;/h2>
&lt;p>At ICLR 2026, authors replied on &lt;strong>73.8%&lt;/strong> of papers. The median reply was &lt;strong>476 words&lt;/strong>, about twice as long as the median review (&lt;strong>248 words&lt;/strong>).&lt;/p>
&lt;p>The right panel of the figure above splits decided papers by what happened during the discussion:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>No author reply at all:&lt;/strong> 5 of 1,209 accepted (&lt;strong>0.4%&lt;/strong>)&lt;/li>
&lt;li>&lt;strong>Authors replied, no reviewer said they raised their score:&lt;/strong> &lt;strong>37.9%&lt;/strong> of 9,865&lt;/li>
&lt;li>&lt;strong>At least one reviewer wrote that they raised their score:&lt;/strong> &lt;strong>61.2%&lt;/strong> of 2,623 (about one in five decided papers)&lt;/li>
&lt;/ul>
&lt;p>Please read these as &lt;strong>associations, not causal effects&lt;/strong>. Authors who expect to be rejected often don&amp;rsquo;t reply at all. &amp;ldquo;Raised my score&amp;rdquo; is detected from comment text, so it misses silent changes. Still, the arithmetic from the previous section explains why the discussion phase matters so much: if your mean is near 5, the rebuttal is where the decision is made.&lt;/p>
&lt;h2 id="6-the-reviews-themselves-are-changing">6. The reviews themselves are changing&lt;/h2>
&lt;p>&lt;img src="fig5_reviews.png" alt="Median review length 2024–2026, and LLM-style words in ICLR reviews">&lt;/p>
&lt;p>&lt;strong>Shorter reviews.&lt;/strong> From 2024 to 2026, ICLR used the same review form (summary, strengths, weaknesses, questions), so lengths can be compared fairly. The number of reviews grew from &lt;strong>28,002 to 75,781&lt;/strong>, while the median length fell from &lt;strong>278 to 249 words&lt;/strong>. More reviews, each a little shorter, is consistent with a stretched reviewer pool, though this data cannot prove that.&lt;/p>
&lt;p>&lt;strong>The rise and fall of &amp;ldquo;delve&amp;rdquo;.&lt;/strong> &lt;a href="https://arxiv.org/abs/2403.07183">Liang et al. (ICML 2024)&lt;/a> showed that LLM-modified text leaves a trail of over-used adjectives and verbs. I tracked four of them in ICLR reviews:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>&amp;ldquo;delve&amp;rdquo;&lt;/strong>: 0.04–0.18% of reviews through 2023, then &lt;strong>1.56% in 2024&lt;/strong> (nine times the 2022 level), 0.33% in 2025, and &lt;strong>0.10% in 2026&lt;/strong>&lt;/li>
&lt;li>&lt;strong>&amp;ldquo;meticulous&amp;rdquo;&lt;/strong>: up &lt;strong>eleven-fold&lt;/strong> over the same 2022→2024 window&lt;/li>
&lt;li>NeurIPS showed the same jump a year earlier: &amp;ldquo;delve&amp;rdquo; went from 0.18% in 2022 to 0.54% in 2023&lt;/li>
&lt;/ul>
&lt;p>The decline after 2024 is the interesting part, and also the easiest to over-read. It could mean reviewers used AI less. It could equally mean that newer models simply stopped saying &amp;ldquo;delve&amp;rdquo;, or that people learned to edit it out. Word counts alone cannot tell these apart.&lt;/p>
&lt;h2 id="three-takeaways-for-authors">Three takeaways for authors&lt;/h2>
&lt;ol>
&lt;li>&lt;strong>Baselines and cost first.&lt;/strong> One in three reviews raises them. Treat comparisons and compute budgets with the same care as the core idea.&lt;/li>
&lt;li>&lt;strong>Near a mean of 5, the rebuttal decides.&lt;/strong> One reviewer moving from 4 to 6 is worth about 25 percentage points. A clear, specific, respectful rebuttal is one of the highest-leverage things you can write.&lt;/li>
&lt;li>&lt;strong>A hot topic is not a shortcut.&lt;/strong> LLM papers are accepted at the average rate. They simply face the most competition.&lt;/li>
&lt;/ol>
&lt;h2 id="methods-and-limitations-briefly">Methods and limitations, briefly&lt;/h2>
&lt;ul>
&lt;li>&lt;strong>Keyword measures are approximate.&lt;/strong> Topic, weakness, &amp;ldquo;score raised&amp;rdquo;, and word counts are regular-expression matches. They miss paraphrases and include some false positives. I read samples of the matches to remove obvious errors; for example, &amp;ldquo;information theory&amp;rdquo; was first being counted as a theory complaint. The exact patterns are listed in the appendix of the &lt;a href="han2026aireviews.pdf">paper&lt;/a>.&lt;/li>
&lt;li>&lt;strong>Final scores.&lt;/strong> OpenReview keeps the latest score, so the score cutoffs mix pre- and post-rebuttal states.&lt;/li>
&lt;li>&lt;strong>Observational.&lt;/strong> Nothing here is a causal estimate. In particular, the rebuttal comparison is confounded by who chooses to reply.&lt;/li>
&lt;li>&lt;strong>Snapshot.&lt;/strong> OpenReview content can be edited or removed. These numbers describe the record as of October 7, 2026.&lt;/li>
&lt;/ul>
&lt;h2 id="a-note-on-how-this-was-made">A note on how this was made&lt;/h2>
&lt;p>I used an AI assistant (Claude) to write the data-collection and analysis code, draw the figures, and draft this post and the paper. Every number was computed directly from the OpenReview data and checked programmatically against the aggregate statistics. I chose the questions and I am responsible for the interpretation. The analysis code is available on request.&lt;/p>
&lt;p>If you work on peer review, meta-science, or reviewer tooling, or you have a hypothesis you&amp;rsquo;d like me to test on this data, I&amp;rsquo;d be glad to hear from you.&lt;/p>
&lt;p>&lt;strong>Download:&lt;/strong> &lt;a href="han2026aireviews.pdf">Working paper (PDF)&lt;/a>&lt;/p></description></item></channel></rss>