logo
|
Blog
    Request Diagnosis

    Vision AI That Distinguishes Subtle Differences Like an Expert

    A Vision AI that distinguishes subtle differences like an expert's eye. See how synthetic data solved degraded nighttime detection performance.
    Pebbly's avatar
    Pebbly
    Sep 16, 2026
    Vision AI That Distinguishes Subtle Differences Like an Expert
    Contents
    What Is Vision AI?How Vision AI Works: How Does It Recognize Images?Vision AI That Distinguishes Wildfire Smoke from Fog: The SphereAX Case1. The Problem: Why Does Vision AI Performance Drop at Night?2. Data Clinic Assessment Results: 10,000 Images on Hand, But a Quality Score of Only 67?3. Pebblous' Solution: Turning Daytime Data into Nighttime Data with a Synthetic Data (“Data Bulk-up”) Strategy4. The Result: A Vision AI with Accurate Recognition Capabilities, Built with Zero Labeling CostsExplore More Vision AI Case Studies

    The images above show wildfire smoke and fog. At first glance, even humans have a hard time telling them apart. Their hazy shapes and blurred backgrounds look remarkably similar. This is especially true when visibility is limited, such as at night or in severe weather.

    So can Vision AI accurately distinguish between the two?

    A well-engineered Vision AI can make that distinction like a seasoned expert. SphereAX, a company developing CCTV-based wildfire detection AI, built an AI that distinguishes nighttime wildfires and smoke with zero labeling costs.

    What Is Vision AI?

    Vision AI is technology that lets machines analyze and interpret the content of images and video. Like humans, it can recognize objects and people, locate them, and detect abnormal situations.

    Today, Vision AI is used across a wide range of industries, including defect inspection in manufacturing, automated logistics sorting, medical imaging analysis, and smart cities. Fire and wildfire detection using CCTV is also a prime example of Vision AI in action.

    How Vision AI Works: How Does It Recognize Images?

    Vision AI derives its results by comprehensively analyzing the color, shape, texture, and surrounding context of objects within an image.

    💡

    Therefore, what ultimately determines the performance of a Vision AI is the quantity and quality of its training data. Only by learning patterns from large volumes of data can the model accurately interpret images it has never seen before.

    However, when the input includes lighting or weather conditions not covered in the training data, the AI's accuracy drops sharply. The "nighttime wildfire detection" challenge that SphereAX faced was a prime example of such a data-scarce domain.

    Vision AI That Distinguishes Wildfire Smoke from Fog: The SphereAX Case

    The long story short: Pebblous' Data Clinic solved SphereAX's nighttime data shortage at zero labeling cost. Here is a detailed look at how the problem was solved.

    What is labeling? It is the process of manually annotating the location and type of objects in an image so that AI can learn from them. Typically, building a new set of 200 nighttime images means paying to film them on-site.

    1. The Problem: Why Does Vision AI Performance Drop at Night?

    "Wildfire smoke is not visible at night."

    "At night, the model cannot distinguish wildfire smoke from fog."

    The model that performed well during the day saw its performance plummet as soon as night fell. SphereAX is a leading AI company in the Daegu region of Korea with strengths in CCTV-based object detection. The company was already running a wildfire and fire detection Vision AI model in production. However, in the process of taking the model to the next level, the company ran into a critical data limitation.

    • The most urgent gap SphereAX needed to close was nighttime data. While the company had secured a sufficient volume of video data from CCTV cameras nationwide, actual nighttime wildfires occur so infrequently that obtaining the necessary data was difficult.

    • In particular, real-world data that could distinguish nighttime wildfire smoke from fog was extremely scarce, making it difficult to collect enough training data to improve model performance.

    The main obstacle to advancing Vision AI is rare data: cases that occur too infrequently to capture at scale.

     Nighttime wildfires, in particular, happen so rarely in the real world that collecting high-quality training data is extremely difficult.

    Even when nighttime data is available, tracing object outlines against dark backgrounds is extremely time-consuming and costly to label. This bottleneck, which could not be solved by data collection alone, was the primary factor slowing down model advancement.

    💡

    This was not a problem that could be solved by simply gathering more data. The real problem was that the necessary data was difficult to obtain in the first place.

    To address this challenge, SphereAX participated in the "Data Quality Assessment and Improvement" program organized by the Daegu Digital Innovation Promotion Agency in Korea. Through this program, Pebblous' Data Clinic began assessing the data quality and analyzing the root cause of the problem.

    2. Data Clinic Assessment Results: 10,000 Images on Hand, But a Quality Score of Only 67?

    First, Pebblous ran a detailed assessment of SphereAX's 10,000 images (5,000 Smoke / 5,000 Fog). Despite the large volume of data, the overall score was 67 points, an "Average" grade.

    Issue Identified

    Details

    Ambiguity between classes

    Wildfire smoke (Smoke) and fog (Fog) are visually very similar, creating a high likelihood of model confusion.

    Excessive duplicate data

    Many images in the Fog class were nearly identical, calling for "Data Diet" deduplication to remove the redundancy.

    Absence of data for specific environments

    Daytime data was sufficient, but data that could guarantee model performance in nighttime and low-light environments was virtually nonexistent.

    In practice, even with 10,000 images, there was almost no training data for the nighttime conditions that actually drove model performance. There was plenty of data, but none of it covered the conditions that mattered most.

    3. Pebblous' Solution: Turning Daytime Data into Nighttime Data with a Synthetic Data (“Data Bulk-up”) Strategy

    The solution Pebblous proposed was not simply to collect more data, but a Data Bulk-up strategy leveraging generative AI.

    The key here was repurposing the existing daytime data. Using daytime images that already contained labeling information (object locations), the approach synthetically transformed the backgrounds into nighttime environments while keeping the labels intact. This made it possible to build nighttime training data without any additional nighttime labeling work.

    Three techniques were applied together:

    • Flux 1 model: A current-generation generative model used for high-quality image synthesis

    • Precision control: Simulating a variety of nighttime lighting conditions by adjusting elements such as the degree of dusk and streetlight effects in stages

    • Concurrent Data Diet: Removing the highly duplicated fog data identified in the earlier assessment, while simultaneously optimizing the dataset by filling that space with scarce nighttime data.

    SphereAX's actual filmed data
    Synthetic data generated by Pebblous based on the actual filmed data

    If you have read this far, however, one question may come to mind.

    "Isn't this just a daytime photo made darker?"

    "Didn't you simply convert existing daytime images into nighttime images and reuse the original labels?"

    That is not the case, for two main reasons.

    1. The data must reflect real nighttime environments as closely as possible.

    Picture an ordinary night. Nighttime is not simply the absence of light. As night falls, a variety of changes occur: moonlight, streetlights, CCTV noise, light bleeding, and degraded sensor quality. In a wildfire detection environment, the variables multiply even further.

    If an AI is trained on images that are merely darkened, it ends up learning "images with a dark filter applied." It never actually learns nighttime conditions. Therefore, to build an AI model that operates reliably in real-world conditions, the training data must reflect these characteristics of the nighttime environment.

    2. The distinctive features of wildfire smoke must be preserved even as the images are converted to nighttime environments.

    What mattered in this project was not simply creating "images that look like night." The data needed to satisfy two conditions at once: a nighttime environment and wildfire smoke.

    🤖

    During generative image transformation, the shape of the wildfire smoke can easily change or its position can shift. In severe cases, the smoke can start to resemble fog, or new elements can appear that were never in the original image.

    In AI training data, however, even small changes can become critical problems. If the position or shape of the wildfire smoke changes during the transformation process, the result is corrupted training data.

    Daytime data turned into nighttime data: Pebblous' synthetic data results

    4. The Result: A Vision AI with Accurate Recognition Capabilities, Built with Zero Labeling Costs

    Pebblous used 200 of SphereAX's original daytime images to generate 200 high-quality nighttime synthetic images (100 fog, 100 wildfire smoke). The clear shapes of the wildfire smoke and fog from the daytime images were preserved intact.

    At the same time, nighttime environments such as night skies, streetlights, and dark backgrounds were rendered so the model could learn from a variety of lighting conditions. This produced training data that enables more accurate distinction between fog and wildfire smoke even in nighttime environments.

    💡

    The biggest change was the reduction in labeling costs. Normally, specialists spend many minutes per image identifying faint objects and labeling them. This project reused the existing daytime labels directly, eliminating a separate labeling process. As a result, labeling costs were reduced to zero. (Estimate based on the equivalent annual labor and resource costs replaced)

    The project also delivered the following benefits:

    • Reduced data collection burden: Replaced the travel and safety costs of directly filming nighttime wildfire scenes, which are both dangerous and rare

    • Shorter AI development timelines: Rapidly built hard-to-obtain nighttime wildfire data, enabling model development without waiting on data collection and shortening deployment time

    The project synthesized several hundred images. After reviewing the completed dataset, SphereAX was pleased enough with the results to ask about a larger follow-up synthesis project.

    Explore More Vision AI Case Studies

    • Manufacturing AI: How do you develop a recycling AI that classifies objects down to their material and gloss? Synthetic data can solve this problem.

    • Defense AI: A defense AI that misidentified friendly forces as enemies is a dangerous AI that defeats its own purpose. See how the problem was solved with data before a dangerous situation could occur.

    Manufacturing AI
    Defense AI

    Should you acquire new data? Or leverage the data you already have?

    The key was not collecting more data, but making better use of the data already on hand. When gathering more data doesn't solve the problem, improving the quality and usability of existing data is often the better starting point for performance gains.

    If data shortages or labeling costs are holding back your Vision AI performance, start by assessing the quality of your current data.

    Data Clinic diagnoses data problems, as in the SphereAX case, and proposes the solutions needed to improve performance, from assessment through synthetic data construction.

    If you would like to receive ongoing AI data quality improvement case studies and practical know-how, subscribe to our newsletter. We regularly deliver real project case studies and the latest news on AI data technology.

    Subscribe to the Data Clinic NewsletterLearn More About Data Clinic, the Data Quality Management Solution
    Share article
    Contents
    What Is Vision AI?How Vision AI Works: How Does It Recognize Images?Vision AI That Distinguishes Wildfire Smoke from Fog: The SphereAX Case1. The Problem: Why Does Vision AI Performance Drop at Night?2. Data Clinic Assessment Results: 10,000 Images on Hand, But a Quality Score of Only 67?3. Pebblous' Solution: Turning Daytime Data into Nighttime Data with a Synthetic Data (“Data Bulk-up”) Strategy4. The Result: A Vision AI with Accurate Recognition Capabilities, Built with Zero Labeling CostsExplore More Vision AI Case Studies

    Pebblous

    RSS·Powered by Inblog