The dataset contains 418 page-level observations across five production sites. Together, those rows recorded 14,195 Google Search impressions and 327 clicks during the 90-day measurement window.
The useful finding is how unevenly the activity was distributed. A total of 383 pages recorded impressions without a click. Home pages represented only 11 rows, but they produced 245 of the 327 clicks. A model trained on the file therefore needs page type, site, and exposure fields kept beside the performance labels. Treating every row as interchangeable would erase the strongest pattern in the sample.
What the dataset joins
Each row combines visible page characteristics with its performance measurements. The input side includes title, description, heading structure, word count, schema types, internal links, canonical status, and template type. The label side includes clicks, impressions, click-through rate, average position, best position, and a first-45-day versus last-45-day split.
Why zero-click rows matter
Rows without clicks are not empty. Every row in this sample recorded at least one impression. They show pages Google exposed but users did not choose during the measured window. That makes them useful negative examples for classification and ranking work, as long as the model also sees the number of impressions and the page type.