The accuracy of PM 2.5 forecasts in Seoul from 2019 to 2023 was assessed using multiple methods. Daily short-term PM 2.5 forecasts, which were provided as four categories ( good , moderate , bad , and very bad ), were directly compared with the corresponding PM 2.5 observation data. Although the probability of detection for days with high PM 2.5 concentrations increased, a simultaneous rise in the false alarm rate resulted in no improvement in the total accuracy and F1-score. To analyze these trends in more detail, the forecast accuracy was further examined based on the PM 2.5 categories. The results showed an annual improvement of 3.65% in the accuracy for the bad category. An analysis based on the announcement time also indicated an increase of over 20% in the accuracy for the bad category for next-day and day-after forecasts. The confusion matrices of forecasted and observed PM 2.5 categories confi rmed this improvement, which was primarily due to a reduction in the number of instances where the moderate category was forecasted as bad . However, the accuracy for the good category showed no signifi cant change and that for the moderate category even declined. These fi ndings highlight the importance of category-specifi c evaluation in air quality forecasting and improving the forecast accuracy, particularly for the good and moderate categories. The reliability of forecasts and their policy relevance may be improved by utilizing these insights and addressing temporal and spatial limitations.