NIST Launches AI Model Evaluation Program to Benchmark Performance on Blind Test Data

National Institute of Standards and Technology, NIST

The National Institute of Standards and Technology launched a new initiative designed to provide standardized, independent evaluations of artificial intelligence models using blind datasets.

    Get the Full Story

    Complete the form to unlock this article and enjoy unlimited free access to all PYMNTS content — no additional logins required.

    yesSubscribe to our daily newsletter, PYMNTS Today.

    By completing this form, you agree to receive marketing communications from PYMNTS and to the sharing of your information with our sponsor, if applicable, in accordance with our Privacy Policy and Terms and Conditions.

    The program is aimed at improving confidence in AI performance while reducing the risk that developers optimize systems for known benchmarks rather than real-world use cases.

    The new Artificial Intelligence Technology Evaluation (AITE) program establishes a sequestered testing environment where AI developers can voluntarily submit models for evaluation against datasets that are not available during model training. By preventing models from being exposed to test data in advance, the program is intended to deliver more rigorous and objective assessments than conventional public benchmarks.

    The launch comes as policymakers and regulators are placing growing emphasis on independent testing of advanced AI systems before deployment. It follows reports that OpenAI models escaped their testing environment and compromised the Hugging Face platform. That incident renewed scrutiny of AI evaluation practices and reinforced calls for testing environments that more closely resemble real-world conditions while limiting opportunities for models to manipulate or memorize evaluation data.

    In an FAQ on the program, NIST said AITE will initially focus on evaluating large vision language models (VLMs) across three scientific and public-interest domains: quantum science, genomics and video-based public safety tasks. Over time, the agency plans to expand the program to include additional subject areas, including natural language processing and other types of AI systems.

    “The infrastructure provided by NIST will provide common data, metrics and scoring to help developers understand the performance of their models,” the agency said in a Monday (July 27) statement announcing the program.

    Unlike many public AI leaderboards, AITE relies on sequestered datasets that remain inaccessible to model developers. NIST said this approach mitigates train-test contamination, in which benchmark data inadvertently becomes incorporated into training datasets, artificially inflating measured performance.

    The program includes two participation tracks.

    Data providers, including researchers and domain experts, may contribute original datasets and associated evaluation tasks. In return, NIST will test leading AI models against those datasets and provide participants with comparative performance measurements.

    Separately, AI developers can submit models for evaluation across the expanding collection of datasets and tasks. Participants will receive detailed information about how their models perform in multiple domains and how they compare with competing systems using identical evaluation criteria, while ensuring that evaluation datasets are not used for model training.

    NIST said testing will begin this summer and expand in four phases. The initial phase will accept only a limited number of external models and evaluation tasks as the agency refines its testing processes with outside collaborators. Later phases will broaden the number of AI applications, datasets, evaluation methods and participating models before transitioning to continuous long-term growth aligned with emerging AI technologies and use cases.

    Participation is voluntary and open to organizations and individuals that agree to AITE’s participation requirements. Any AI model capable of complying with the program’s API and evaluation criteria may be submitted for testing.

    NIST cautioned against overinterpreting early results. Because the initial version of AITE includes a relatively small number of datasets and evaluation tasks, the agency said performance on those benchmarks should not be assumed to predict performance across all real-world applications. As the repository of datasets grows, broader conclusions about model capabilities may become more reliable. Publication of evaluation results should not be construed as government approval or endorsement of any commercial AI product.

    For regulators and industry, the initiative represents another step toward establishing common methodologies for measuring AI systems. By combining standardized metrics, blind testing and publicly available results, the program seeks to provide a more credible basis for comparing models as governments continue to develop policies governing AI safety, reliability and deployment.

    For all PYMNTS AI coverage, subscribe to the daily AI Newsletter.