Guide

Stripping EXIF Is No Longer Enough

By the Smol team11 min read

Removing EXIF deletes a photo’s claim about where it was taken. It does not touch the evidence in the picture. Systems that infer location from pixels alone now place roughly one in four cue-rich outdoor photographs within a kilometre of the right spot, with no metadata at all, and that figure comes from a peer-reviewed measurement rather than a vendor.

Which means the old advice is now half an answer. Strip the metadata, absolutely. Just stop treating it as the finish line.

This article sticks to numbers with a citation and a date. Where the evidence is a preprint rather than peer-reviewed, it says so. Where there is no evidence at all, it says that too, because one of the most useful findings here is a gap.

What can image-geolocation models actually do?

The field reports accuracy as the share of test images placed within a set of distance bands. The coarse four come from Hays and Efros’ IM2GPS paper at CVPR 2008; a 1 km street-level band was added later and is now standard.

SystemTest set1 km25 km200 kmEvidence
PIGEON (Stanford)Street View holdout5.4%40.4%78.3%CVPR 2024, peer-reviewed
PIGEOTTO (same authors, general photos)Im2GPS3k11.3%36.7%53.8%CVPR 2024, peer-reviewed
GeoCLIPIm2GPS3k14.1%34.5%50.7%NeurIPS 2023, peer-reviewed
GPT-4o, zero-shotIm2GPS3k14.4%38.9%55.8%AAAI-26, peer-reviewed
GeoSpy (commercial)GeolocationHub26.5%51.1%85.7%PoPETs 2025, peer-reviewed

That last row is the one to hold onto, with its caveat attached. The 26.5% figure is from Liu et al., Proceedings on Privacy Enhancing Technologies 2025(4), across 20,000 images. The dataset deliberately excludes indoor shots, blurry, underexposed and overexposed frames, tunnels, ceilings and close-up textures. So 26.5% is an upper bound on easy photographs, not an average over your camera roll.

PIGEON’s authors are worth reading on their own work rather than being summarised. On GeoGuessr, their model was placed in the Champion Division, the top 0.01% of players, across 458 matches. Their conclusion in the paper:

“While a major limitation of today’s image geolocalization technologies (including ours) is that they are unable to make street-level predictions reliably, researchers ought to carefully consider the risk of potential misuse of their work as such technologies get increasingly precise.”

Haas, Skreta, Alberti and Finn, PIGEON: Predicting Image Geolocations, CVPR 2024

Both halves of that sentence are load-bearing. Street level is still unreliable. City level is not.

Is this just the model reading the EXIF?

No, and this has been tested directly rather than assumed.

Bellingcat ran 500 tests across 20 models in June 2025, using their own photographs that had never been published online. Their methodology sentence is the cleanest control in the literature: “Each LLM was given a photo that had not been published online and contained no metadata.” The models still worked, and three of them beat Google Lens.

A 2025 preprint, GEO-Detective, ran the ablation the other way round: it modified the EXIF and measured what changed. Country-level accuracy moved from 50.0% to 52.0%, which is noise. The authors’ reading is that the system “does not rely on metadata during reasoning.”

There is a smaller, funnier data point from the April 2025 wave of viral geolocation posts. Sam Patterson ran a controlled GeoGuessr match between OpenAI’s o3 and a ranked human player and published the replayable challenge. He also found that uploading a file to ChatGPT through a browser strips its EXIF before the model ever sees it, so he had to zip the files to get metadata in at all. When he fed the model deliberately spoofed coordinates, it told him it was not buying them.

Which settles the point mechanically. For the entire “it is just reading the EXIF” argument, the metadata was not there.

The vendors are explicit about this too, because it is their selling point. Graylark’s product page for Raven says it “reads pixels, identifies signals, and returns actionable leads in seconds, with no metadata required.”

What are these systems bad at?

A lot, and the failures are as well documented as the successes. This is where the non-sensational version of the argument earns its keep.

  • Uniformly sampled locations. On GWS15k, a benchmark built by sampling Street View evenly rather than from photogenic places, the best reported result places 0.7% of images within 1 km. Curated benchmarks flatter these models badly.
  • Indoors. Measured by scene type at ECCV 2018: 14.3% of indoor images landed within 25 km against 36.3% of urban ones. Natural landscapes were worst at street level, 3.3% within 1 km.
  • Anywhere with thin Street View coverage. A 2025 preprint measured Gemini 2.5 Pro at 6.4% within 1 km on Chinese-region footage against 65.6% on European and American footage. Same model, same task, a tenfold gap that follows the training data.
  • Their own list. PIGEON’s appendix names “tunnels, bodies of water, poorly illuminated areas, forests, indoor areas, and soccer stadiums” as its hardest cases.

And a genuinely counterintuitive result from that same appendix, which is the reason the obvious advice does not always hold. The authors report their general-photo model working “surprisingly well in indoor scenarios, even when images are blurred.” It placed a close-up photograph of leaves within 671 km from the flora alone, and a person drinking from a red cup within 13 km.

Thirteen kilometres from a red cup is not a street address. It is also not nothing, and it came from a photo with no sky, no signage and no horizon.

How fast is this actually moving?

Fast between 2016 and 2024, and roughly flat since. Both halves of that matter, and the second half almost never gets written down.

The first serious planet-scale system was Google’s PlaNet, published at ECCV 2016. Its authors pitted it against ten well-travelled human subjects in GeoGuessr and reported the result plainly: “PlaNet won 28 of the 50 rounds with a median localization error of 1131.7 km, while the median human localization error was 2320.75 km.” Players were allowed to pan and zoom, just not to walk to the next panorama. The whole model was 377 MB, which the authors noted would fit on a phone.

Eight years later PIGEON reported a median error of 44.35 km on a 5,000-location Street View holdout. Different test set, so it is not a clean like-for-like, but the direction and the order of magnitude are not in doubt: a median error measured in thousands of kilometres became one measured in tens.

Then it stopped. A community benchmark of 100 Street View stills, which are structurally free of metadata, scored GPT-5 in August 2025 at 3,937 points against o3-high’s 3,929 from four months earlier. Gemini 2.5 Pro led both at 4,085, which is its own correction to the coverage of that period: the model everyone wrote about was never the best one. Bellingcat’s August 2025 retest of 24 models found GPT-5 had regressed against the previous generation on their geolocation set.

So the sensible expectation is not a smooth curve toward street addresses. It is a capability that arrived quickly, plateaued at city level, and is now limited by how much of the world has been photographed from the road.

Who is running this outside a research lab?

The visible commercial case is Graylark Technologies. Its consumer-facing product GeoSpy was the subject of 404 Media reporting by Joseph Cox on 20 January 2025, after which the public tier was closed. Graylark renamed the platform Raven in April 2026, stating on its own site: “In April 2026, GeoSpy became Raven, Graylark’s frontline visual intelligence platform for law enforcement and government.” 404 Media reported confirmed law-enforcement purchases in February 2026, naming the Miami-Dade Sheriff’s Office and the LAPD. The company announced a $10.7 million seed round on 29 August 2026.

There is a gap between what the vendor claims and what has been measured, and it is large enough to be the most useful thing on this page. As of 25 September 2026, geospy.ai still advertises:

“Delivering up to meter level accuracy, state of the art computer vision models all in an easy to use interface.”

geospy.ai, retrieved 25 September 2026. Vendor claim, not a measurement.

The vendor publishes no hit rate, no test set and no error distribution anywhere. The only peer-reviewed measurement of the same product puts it at 26.5% within a kilometre on deliberately favourable images, with a mean error of 545.8 km. “Up to metre level” and “one in four within a kilometre” are both true statements about the same system. Only one of them is a description of what it does.

Treat any figure you see on this subject the same way. A number without a test set is marketing.

What actually reduces the risk?

Ranked by the strength of the evidence behind them, not by how good they sound.

MitigationMeasured effectEvidence
Mask visible text: signage, storefronts, plates, menuso3’s median error rose from 3.7 km to 186.9 km. City-level accuracy fell from 46.6% to 24.0%.IMAGEO-Bench, 2025 preprint
Capture less of the sceneDropping a four-image panorama for a single frame took PIGEON’s median error from 60.8 km to 131.1 km and country accuracy from 87.6% to 74.7%.CVPR 2024, peer-reviewed. Note: this is about framing at capture, not cropping afterwards.
Blur the backgroundRoughly half of clean accuracy retained in one narrow regional-classification test. No published measurement on global geolocation.2026 preprint, weak
Reduce resolutionNo published measurement exists.See below
Remove EXIFCountry accuracy moved 50.0% to 52.0% when metadata was altered.GEO-Detective, 2025 preprint

The resolution row is not an oversight. We looked for a study measuring how much downscaling degrades planet-scale geolocation and did not find one, for any model, in any year. What adjacent peer-reviewed work exists points the wrong way: a CVPR 2022 benchmark concluded that for geo-localization “using the highest available resolution is in most cases superfluous, and often even detrimental,” and an ECCV 2024 robustness study found the effects of JPEG compression and pixelation “relatively minor.” A 2026 preprint measured GPT-5 at 20.0% within 1 km on images downscaled to 1,080 px and saved at JPEG quality 85, although that run had no untransformed control row and so cannot isolate the effect.

The honest statement is therefore: nobody has measured it, and the surrounding evidence suggests it helps very little. If you see a page telling you to resize your photos for privacy, it is not citing anything, because there is nothing to cite.

Cropping has a specific trap of its own, which we measured directly. Take a 6,241 × 4,161 photograph, crop it to 2,000 pixels wide, and carry the metadata across the way most conversion workflows do. The embedded EXIF thumbnail in the output is still 256 × 171 and byte-identical to the one in the uncropped original. You cropped the picture and shipped a small copy of what you cropped out. Cropping only works if the metadata goes too.

Two mitigations that work in the lab and are not available to you. Adversarial perturbation, which cut GPT-4o’s street-level accuracy from 7.3% to 1.1% in an AAAI-26 paper, requires running an optimizer against surrogate models. Visible watermarking works by triggering the model’s refusal behaviour rather than removing information, which an unaligned model or a crop defeats. Neither ships as a user-facing tool today.

So should you still bother removing EXIF?

Yes. The argument for it just changed shape.

EXIF GPS is exact, free to read, and requires no model. A pair of coordinates and an altitude reading identifies a floor of a building. The best published inference result on favourable images is a one-in-four chance of landing within a kilometre. Those are not comparable threats, and removing the first one costs you four seconds.

What has changed is that metadata removal is now the floor rather than the solution. The useful mental model is two independent questions:

  • What does the file say? Answer: nothing, once you have stripped it. This is solved, cheaply and completely.
  • What does the picture show? Answer: whatever you pointed the camera at. This is an editorial decision, not a technical one, and no tool makes it for you.

For most photographs, the second question does not matter. A close-up of a meal, a product on a white background, a screenshot: these are not located. For a photograph of the street outside your house, no amount of metadata hygiene helps, and the only real mitigation is to not post that photograph.

It is also worth knowing that the regulatory machinery does not currently reach this. Privacy International’s February 2026 report “Nowhere to Hide?” argues that these deployments sit in a structural gap: the rules are built around biometric data, meaning features of a person’s body, while image geolocation reads the background. Nothing about your legal position changes because you stripped a file.

When Smol is not the answer

Nothing in this article is a product problem, and we are not going to pretend otherwise. We make Smol, a Mac app that removes metadata from a folder of images without re-encoding them. It solves the first of the two questions above completely and the second one not at all.

For a single photo, exiftool -all= photo.jpg is free and produces a byte-identical result to ours, which we verified. For a folder, one exiftool process pointed at a directory cleared 50 files in 465 ms against our 5,281 ms. Smol is $29 once and what it buys you is a drop target, a Finder Quick Action and a watched folder rather than a faster stripper.

And if what you actually want is to be un-findable in a photograph, no metadata tool on the market delivers that, including ours. The mitigations with real evidence behind them are compositional: do not include readable signage, do not include the front of the building, do not include the view from the window. That is a decision you make before you press the shutter.

The mechanics of the stripping itself, with every route measured, are in how to remove EXIF data on Mac. The location-specific version, including the Photos library that keeps its own copy of your coordinates, is in stripping GPS from photos before you share them. If you want to know what the tags are before you delete them, what EXIF data actually contains walks a real dump line by line, and our help article is the three-step version. For documents rather than photos, the equivalent question is answered in compressing a PDF on Mac.

Frequently asked questions

Can AI find the location of a photo without EXIF data?

Yes, within limits that are now measured. A peer-reviewed 2025 study in Proceedings on Privacy Enhancing Technologies put one commercial system at 26.5% of images within 1 km and 51.1% within 25 km, on a dataset that excluded indoor, blurry and low-detail photographs. Research models on general photo benchmarks land in the 11 to 14% range at 1 km. None of them needs metadata.

Does removing EXIF data still protect my privacy?

It removes an exact, free, machine-readable statement of where you stood, which is worth doing and takes seconds. It does not make the photograph location-anonymous. Treat metadata removal as necessary rather than sufficient, and make a separate decision about what the picture itself shows.

Does resizing or compressing a photo stop AI geolocation?

There is no published measurement of this, for any model, which is itself worth knowing. The adjacent peer-reviewed evidence points the wrong way: a CVPR 2022 benchmark found the highest resolution is often unnecessary for geo-localization, and an ECCV 2024 robustness study found JPEG compression and pixelation had relatively minor effects. Assume downscaling helps very little.

What is GeoSpy and is it still available?

GeoSpy was an image-geolocation product from Graylark Technologies. Its public tier closed after 404 Media reporting in January 2025, and the company renamed the platform Raven in April 2026, describing it on its own site as a visual intelligence platform for law enforcement and government. 404 Media reported confirmed purchases by the Miami-Dade Sheriff’s Office and the LAPD in February 2026.

What actually reduces how well a model can geolocate my photo?

Masking visible text is the strongest measured mitigation: covering signage, storefront names and plates raised one model’s median error from 3.7 km to 186.9 km in a 2025 preprint. Capturing less of the scene also measurably hurts these models. Blurring has one weak published result, and resolution reduction has none at all.

Are these models accurate everywhere in the world?

No, and the gap is large. A 2025 preprint measured the same model at 6.4% within 1 km on Chinese-region imagery against 65.6% on European and American imagery. Accuracy tracks Street View coverage and dataset composition, so results reported on Western benchmarks do not transfer.

Keep reading