Model VAPT

Model & Weights Pen Testing

VAPT for AI/ML models, custom fine-tuned models, and model weights

Continuous pen testing for AI/ML models as assets — model extraction, inversion, membership inference, backdoor detection, adversarial robustness, unsafe-serialisation / loader RCE, model-registry supply chain, and serving-infrastructure hardening. Compliance mapping for ISO 42001, NIST AI RMF, EU AI Act, SOC 2, GDPR, HIPAA, and DPDP Act.

How It Works

Four steps. One continuous pen testing loop.

1

Model & Asset Mapping

Catalogue the models in scope — foundation checkpoints, custom fine-tuned derivatives, distilled student models, quantised variants — plus their storage locations, serving infrastructure, access-control boundaries, and the registries or hubs they move through.

2

Attack-Surface Pen Testing

Pen test against model-specific attack classes — extraction via query enumeration, inversion to reconstruct training data, membership inference, backdoor detection, adversarial-example robustness, and the data-poisoning paths that could compromise future training runs.

3

Weights, Format & Supply Chain

Validate how model weights are stored, loaded, and shipped — unsafe pickle in legacy checkpoints, safetensors / GGUF / ONNX integrity, signed-release provenance, registry trust (HuggingFace, Vertex Model Registry, SageMaker Model Registry, custom OCI registries), and the loader code paths that execute untrusted bytes.

4

Serving & Evidence

Test the serving layer — Triton, vLLM, TGI, Ray Serve, SageMaker endpoints, Vertex AI, custom FastAPI wrappers — for authentication, rate limiting, inference-time injection, and side-channel leakage. Reports map to ISO 42001, NIST AI RMF, EU AI Act, SOC 2, and ISO 27001 Annex A.

Key Features

Full model attack surface — weights, loader, registry, serving

Model Extraction (Black-Box)

Query-based extraction via systematic probing of the inference endpoint — reconstruct an approximate copy of a proprietary model from public API access

Model Extraction (Grey-Box)

Where the attacker knows architecture or has partial access — accelerated extraction via gradient estimation or embedding leakage

Model Inversion

Recover approximate training examples from the deployed model's outputs or (where exposed) its weights — a privacy concern for models trained on sensitive data

Membership Inference

Determine whether a specific record was in the training set — a GDPR / HIPAA / DPDP concern for models trained on regulated personal data

Backdoor / Trojan Detection

Test models supplied from untrusted sources or trained on partially-attacker-controlled data for trigger-based backdoors

Adversarial Robustness Testing

Validate model behaviour under adversarial perturbations — FGSM, PGD, Carlini-Wagner for vision; TextFooler, BERT-Attack, universal triggers for NLP

Unsafe Pickle / Loader RCE

Validate every model-loading code path in training, fine-tuning, and inference pipelines for pickle acceptance and the arbitrary-RCE class that follows

Safetensors / GGUF / ONNX Integrity

Where safer formats are claimed, verify adoption, loader-side validation, and absence of fallback-to-pickle paths

Registry & Supply Chain

Validate trust in HuggingFace Hub, Vertex Model Registry, SageMaker Model Registry, Azure ML Model Registry, custom OCI registries — signing, provenance, access control, and malicious-model detection

Weight Encryption at Rest

For high-IP model weights, validate encryption at rest, access control, exfiltration monitoring, and insider-threat resistance across training storage, model registry, and serving infrastructure

Fine-Tuning Data Leakage

Test whether the fine-tuning dataset — often containing privileged, PII, or ePHI content — can be recovered through model probing, including recent memorisation-attack techniques

Serving-Layer Pen Testing

Full attack-surface testing of Triton, vLLM, TGI, Ray Serve, KServe, SageMaker, Vertex AI, Azure OpenAI, OpenAI-compatible proxies, custom FastAPI serving — auth, rate limits, cost DoS, side channels

Quantisation / Distillation Integrity

Validate that quantised and distilled derivatives of a trained model do not degrade safety or policy behaviour in ways attackers can exploit

Watermarking Integrity

For models carrying output watermarks (invisible attribution or policy-provenance markers), test watermark robustness and attacker-side removal

Multi-Model / Mixture Routing

In mixture-of-experts or router-dispatched deployments, test whether routing can be manipulated to force queries into weaker or policy-exempt experts

Benefits

Why teams choose TigerStrike for their security needs

Weights-as-Crown-Jewel Testing

For organisations where fine-tuned model weights are intellectual property — the trained parameters encode months of training spend and sometimes privileged data. We treat weights with the same rigour pen testing applies to a source-code repository: access control, encryption at rest, exfiltration paths, and insider-threat resistance.

Weights-as-Crown-Jewel Testing

Model Extraction Resistance

Test whether a competitor or attacker can reconstruct an approximate copy of your model through systematic query enumeration against your inference endpoint. Includes black-box and grey-box extraction techniques, defences via rate limiting, output noise, and query-pattern anomaly detection.

Model Extraction Resistance

Training-Data Privacy Attacks

Model inversion attacks recover approximate training examples from a trained model's outputs or weights. Membership-inference attacks determine whether a specific record was in the training set. Both matter when training data includes PII, ePHI, PCI data, or legally privileged content — we test what your deployed model actually leaks.

Training-Data Privacy Attacks

Backdoor & Trojan Detection

Models trained on partially-untrusted data or sourced from third-party registries can carry trigger-based backdoors — benign on normal inputs, adversarial on inputs containing a specific trigger. We test for known backdoor patterns and validate your detection controls on supplied test models.

Backdoor & Trojan Detection

Unsafe Serialisation & Loader RCE

Legacy model formats — pickle, torch .pt / .pth, joblib — execute arbitrary code on load. HuggingFace and other registries have hosted malicious pickle checkpoints in the wild. We validate the loader code paths in your training and inference pipelines, test for unsafe pickle acceptance, and verify safetensors / GGUF / ONNX adoption where it matters.

Unsafe Serialisation & Loader RCE

Serving Infrastructure Hardening

Test the serving layer end-to-end — Triton Inference Server, vLLM, TGI, Ray Serve, KServe, SageMaker endpoints, Vertex AI, Azure OpenAI, OpenAI-compatible proxies, custom FastAPI wrappers. Covers authentication, rate limiting, prompt-injection via serving-layer features, cost DoS, and side-channel leakage via timing or token-probability.

Serving Infrastructure Hardening

Frequently Asked Questions

Ready to get started?

Start securing your applications today with TigerStrike's AI-powered penetration testing platform.

Book a Demo