find-the-original-image
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseFind the original image
查找图片原始来源
The goal is almost never "find a match". It is find the earliest publication
and read its page. A match tells you the image exists elsewhere; the earliest
page tells you the photographer, the date, the caption, and the names — which is
what you actually pivot on.
The beginner mistake: uploading the full frame to one engine, getting nothing,
and concluding the image is unindexed. Cropping to one distinctive object and
re-searching finds things full-frame search cannot.
我们的目标几乎从来不是“找到匹配结果”,而是找到最早的发布页面并阅读其内容。匹配结果只能告诉你该图片在其他地方存在;而最早的页面能提供摄影师信息、拍摄日期、图片说明以及相关人物姓名——这些才是你真正需要的关键线索。
新手常犯的错误:将整张图片上传至一个搜索引擎,没有得到结果就断定该图片未被索引。其实,裁剪出图片中一个独特的元素再重新搜索,往往能找到整张图片搜索无法发现的内容。
Pick your engine by what you are holding
根据持有内容选择合适的引擎
| You have | Start with | Why |
|---|---|---|
| A face | Yandex | Its index is built around facial similarity, so it returns different people who look alike and the same person in other photographs. No other general engine does this. |
| A face, and Yandex fails | A dedicated face engine (see below) | Only after you have cleared the legal and consent questions. |
| A street scene outside North America / Western Europe | Yandex | Deeply indexed Russian, Central Asian, Eastern European, Turkish and Chinese web content that Google under-crawls. |
| A product, book cover, artwork, plant, animal | Google Lens | Object and entity recognition, tied to Shopping and Knowledge Graph. |
| Text inside the image | Google Lens | It OCRs the frame and lets you search the extracted string. Often the text is the answer and the image search is irrelevant. |
| A landmark or a well-known building | Google Lens | Landmark classification is its strongest single feature. |
| A specific region of a cluttered photo | Bing Visual Search | Draw a box on the uploaded image and it re-searches only that region — the fastest crop-and-retry loop of any engine. |
| A press photo, meme, or anything you suspect is old | TinEye | The only major engine that sorts by oldest and that reliably surfaces modified copies. |
| Chinese-language or China-hosted content | Baidu image search | Coverage the others simply do not have. |
Run at least three. They disagree constantly, and that disagreement is
information: TinEye finding an exact copy from years back while Lens finds only
recent reposts is the signature of recycled media.
| 你拥有的内容 | 首选引擎 | 原因 |
|---|---|---|
| 人脸图片 | Yandex | 其索引基于面部相似度构建,不仅能返回同一人物的其他照片,还会返回外貌相似的不同人物。其他通用搜索引擎都不具备这一功能。 |
| 人脸图片,但Yandex搜索失败 | 专用人脸搜索引擎(见下文) | 仅在你确认符合法律规定及相关授权后使用。 |
| 北美/西欧以外的街景图片 | Yandex | 深度索引了俄罗斯、中亚、东欧、土耳其和中国的网络内容,这些是Google爬虫覆盖不足的区域。 |
| 产品、书籍封面、艺术品、植物、动物图片 | Google Lens | 具备物体和实体识别功能,与购物服务及知识图谱关联。 |
| 图片中包含文字 | Google Lens | 可对图片进行OCR识别,并允许你搜索提取出的文字。通常文字本身就是答案,图片搜索反而无关紧要。 |
| 地标或知名建筑图片 | Google Lens | 地标分类是其最突出的功能。 |
| 杂乱照片中的特定区域 | Bing Visual Search | 可在上传的图片上绘制选区,仅针对该区域重新搜索——这是所有引擎中最快的裁剪重试流程。 |
| 新闻照片、表情包或疑似旧图 | TinEye | 唯一能按最早发布时间排序,并可靠识别修改后副本的主流引擎。 |
| 中文内容或中国境内托管的内容 | 百度图片搜索 | 拥有其他引擎不具备的覆盖范围。 |
至少使用三个引擎进行搜索。它们的结果经常不一致,而这种不一致本身就是信息:比如TinEye找到了多年前的精确副本,而Lens只找到近期的转发内容,这就表明该媒体是被循环使用的。
What each engine is actually doing
各引擎的实际工作原理
Yandex matches on visual similarity with a strong face component. Its results
page groups "sites containing this image" separately from "similar images" —
only the first group is evidence. It tolerates crops, rotation and heavy
recompression better than the others.
Google Lens has moved away from whole-image duplicate matching toward "what
is this, and what can I sell you". For provenance work use the option that lists
pages containing the image rather than the visual-match carousel, and expect it
to return visually similar but unrelated photos as if they were matches.
Bing Visual Search sits between the two. Its region-select tool is the reason
to use it: no download, crop, re-upload cycle.
TinEye is crawl-based and comparatively small — plenty of images return zero
results, and absence from TinEye proves nothing. What it does that nothing else
does: exact and near-duplicate matching with the ability to sort by oldest, and
detection of copies that have been cropped, colour-shifted or watermarked, which
it will show you side by side against your input.
Yandex:基于视觉相似度匹配,面部识别能力较强。其结果页面会将“包含该图片的网站”与“相似图片”分开分组——只有第一组可作为有效证据。相比其他引擎,它更能容忍裁剪、旋转和重度重新压缩的图片。
Google Lens:已从整张图片的重复匹配转向“识别物体并提供相关推荐”。对于溯源工作,应选择列出包含该图片的页面的选项,而非视觉匹配轮播图,并且要注意它可能会返回视觉相似但无关的照片作为匹配结果。
Bing Visual Search:介于前两者之间。其区域选择工具是使用它的核心原因:无需下载、裁剪、重新上传的循环操作。
TinEye:基于爬虫构建,索引规模相对较小——很多图片搜索会返回零结果,因此未在TinEye中找到结果并不能证明任何事情。它的独特之处在于:精确和近似重复匹配,支持按最早发布时间排序,还能检测经过裁剪、调色或添加水印的副本,并将这些副本与你的输入图片并排展示。
Method
操作方法
- Fix the input. Get the highest-resolution copy you can — the file itself,
not a screenshot of it. Screenshots add a resize and a recompression that cost
you matches. Strip nothing yet; run on the original before you start editing copies.
secrets-in-file-metadata - Full frame, all engines. Cheap, sometimes instant.
- Crop and re-search. This is the highest-yield step in the whole skill. A full frame's fingerprint is dominated by the background; crop tightly to the face, the sign, the logo, the vehicle, the tattoo, the building corner, and search each crop separately. Reposts get cropped and re-framed, so the crop often matches when the frame does not.
- Flip horizontally. Mirroring is a routine way to dodge automated matching and a routine artifact of screen-recording and video reposting. Flip and re-run the same engines.
- Preprocess and retry on low-quality inputs — upscale, denoise, correct levels. See reference/preprocessing.md.
- For video, extract keyframes and search them as stills. Search a frame from the start, the middle and the end, plus any frame containing a sign or a face. A keyframe hit is usually what breaks a video case.
- Sort for oldest. On TinEye, sort by oldest. Then treat that date as a ceiling, not an answer, and go to step 8.
- Read the match pages. Open them. Harvest: photographer credit, agency,
caption, named people, on-page date, filename in the image URL (often
), and any surrounding article text. This is where the selectors are.
2019-04-city-event-03.jpg
- 优化输入图片:获取最高分辨率的图片副本——最好是文件本身,而非截图。截图会带来缩放和重新压缩,导致匹配成功率下降。在开始编辑副本前,先对原始文件运行工具检查元数据。
secrets-in-file-metadata - 全图搜索,覆盖所有引擎:操作简单,有时能快速得到结果。
- 裁剪后重新搜索:这是整个技能中产出最高的步骤。整张图片的指纹会被背景主导;裁剪出紧密的人脸、标识、标志、车辆、纹身、建筑角落等区域,分别搜索每个裁剪后的图片。转发内容通常会被裁剪和重新构图,因此裁剪后的图片往往能匹配成功,而整张图片却不行。
- 水平翻转图片:镜像处理是规避自动匹配的常用手段,也是录屏和视频转发时的常见产物。翻转图片后重新使用相同引擎搜索。
- 预处理后重试:针对低质量输入图片——进行放大、降噪、校正色阶等操作。详见reference/preprocessing.md。
- 针对视频:提取关键帧并将其作为静态图片搜索。搜索视频开头、中间、结尾的帧,以及包含标识或人脸的任何帧。关键帧匹配通常是破解视频案例的突破口。
- 按最早发布时间排序:在TinEye中,按最早发布时间排序。然后将该日期视为上限,而非最终答案,进入步骤8。
- 阅读匹配页面:打开匹配页面,收集以下信息:摄影师署名、代理机构、图片说明、提及的人物姓名、页面显示日期、图片URL中的文件名(通常为格式),以及周围的文章文本。这些就是你需要的关键线索。
2019-04-city-event-03.jpg
Establishing first publication, not just a match
确定首次发布来源,而非仅找到匹配结果
An earliest-known-copy date is a claim about your search coverage, not about the
world. To harden it:
- Check the match page's own date against — the archive's first capture of the URL bounds when the page really existed, and the page's displayed date can be back- or forward-dated by its CMS.
read-deleted-pages - Use date-restricted search operators via to look for earlier text mentions of the same event or caption.
google-like-a-spy - Look for the same image at a larger resolution. The largest version is usually closest to the source; a wire agency's copy will be bigger than a Twitter repost of it. TinEye's biggest-image sort is useful here.
- Check the obvious stock libraries. A photo credited to three different news events in three countries is stock, and the "original" is a licensing page.
- Watermarks and agency slugs in the frame beat any index. Crop the watermark and search that alone.
已知最早副本的日期只是关于你的搜索覆盖范围的结论,而非客观事实。要强化这一结论:
- 使用工具,将匹配页面自身显示的日期与存档首次捕获该URL的日期进行对比——存档首次捕获的日期决定了页面实际存在的时间,而页面显示的日期可能被其CMS系统篡改。
read-deleted-pages - 通过工具使用日期限制搜索运算符,查找同一事件或图片说明的更早文本提及记录。
google-like-a-spy - 寻找同一图片的更高分辨率版本。最大尺寸的版本通常最接近原始来源;通讯社的副本会比Twitter上的转发版本更大。TinEye的按最大图片排序功能在此处很有用。
- 检查知名图库素材网站。一张被三个国家的三个不同新闻事件引用的照片很可能是图库素材,其“原始来源”就是授权页面。
- 图片中的水印和代理机构标识比任何索引都更可靠。裁剪出水印单独搜索。
Face-specific engines and when not to use them
专用人脸搜索引擎及使用限制
PimEyes and FaceCheck.ID crawl the open web for faces and match on biometric
similarity. They find people that no general engine will. They also carry real
exposure:
- A facial template is biometric data under GDPR Article 9 — a special category needing an explicit lawful basis, not just a legitimate interest. Illinois BIPA and Texas CUBI create private rights of action and statutory damages for collecting biometric identifiers without written consent.
- PimEyes' own terms position it as a tool to search for your own face and offer an opt-out. Running a third party's face through it is against those terms regardless of your intent.
- Search4Faces indexes faces from Russian social platforms and is the practical option when a subject's footprint is VK or Odnoklassniki.
Use face search when you have a documented authorization or a legitimate
protective purpose — verifying a counterparty in a fraud case, identity
verification with the subject's consent, missing-persons work, or checking your
own exposure. Do not use it to identify a stranger from a photo, to attach a name
to a face in a protest crowd, or to locate a private individual. Note in the case
file that you ran it and why. And treat a face-engine hit as unconfirmed on
its own — look-alike false positives are common and the engine gives you no
reasoning to audit.
PimEyes和FaceCheck.ID会在公开网络上爬取人脸信息,并基于生物特征相似度进行匹配。它们能找到通用搜索引擎无法发现的人物信息,但也存在实际风险:
- 根据GDPR第9条,面部模板属于生物特征数据——这是一个特殊类别,需要明确的合法依据,而非仅仅是合法利益。伊利诺伊州的BIPA法案和德克萨斯州的CUBI法案规定,未经书面同意收集生物特征标识符会引发私人诉讼,并需支付法定赔偿金。
- PimEyes的服务条款将其定位为搜索你自己人脸的工具,并提供退出选项。无论你的意图如何,使用它搜索第三方人脸都违反其条款。
- Search4Faces索引了俄罗斯社交平台的人脸信息,当调查对象的足迹在VK或Odnoklassniki时,它是实用的选择。
仅在你有书面授权或合法保护目的时使用人脸搜索——比如验证诈骗案件中的交易对手、经当事人同意的身份验证、失踪人员调查,或检查你自己的信息暴露情况。不要用它识别陌生人的照片、给抗议人群中的人脸标注姓名,或定位私人个体。在案件文件中记录你使用了该工具及原因。并且,将人脸引擎的匹配结果视为未确认——外貌相似的误判很常见,且引擎不会提供可审核的推理过程。
Where this goes wrong
常见误区
- Similar is not same. Every engine mixes "pages with this image" and "images that look like this" into one visual grid. Only the first is evidence. Confirm a candidate by opening it and comparing pixel-level detail — a hand position, a fold in cloth, a reflection.
- Absence is the default, not a finding. Non-indexed, freshly published, login-walled, and app-only images return nothing. "No results across four engines" means the image is not in those indexes, which is weak evidence for originality and no evidence at all of authenticity.
- Stock photography poisons attribution. A "CEO" headshot that appears on forty unrelated sites is a stock model, and the entity using it is probably fake — that itself is a finding, and a good one.
- AI-generated images have no origin to find. Zero matches plus generative
tells is a distinct outcome; hand it to rather than concluding the photo is an unpublished original.
is-this-photo-real - Engines rewrite history. Results are not reproducible: indexes drop pages, results reorder, and the hit you found last week may be gone. Screenshot the results page and archive the match URL the moment you find it.
- Crops mislead in the other direction. A crop of grass or sky returns thousands of confident, meaningless matches. Crop to what is unusual, not merely present.
- You may be leaking. Uploading a client's image to a commercial engine
discloses it to that vendor and may deanonymise your interest. For sensitive
material, search a crop that omits identifying detail — or don't search. See
.
investigate-without-getting-made
- 相似不等于相同:每个引擎都会将“包含该图片的页面”和“相似图片”混合到同一个视觉网格中。只有前者可作为证据。通过打开候选页面并对比像素级细节(比如手部姿势、衣物褶皱、反光)来确认是否为同一图片。
- 无结果是常态,而非结论:未被索引、刚发布、需要登录或仅在应用内可见的图片会返回无结果。“四个引擎都无结果”仅意味着该图片不在这些引擎的索引中,这不能作为图片原创性的有力证据,更不能证明其真实性。
- 图库素材会干扰归因:一张出现在四十个无关网站上的“CEO”头像很可能是图库模特照片,使用该图片的实体很可能是虚假的——这本身就是一个重要发现。
- AI生成的图片没有原始来源:无匹配结果加上AI生成特征是一种明确的情况;应将其交给工具处理,而非断定该照片是未发布的原创内容。
is-this-photo-real - 引擎会改写历史:结果不可重现:索引会删除页面,结果会重新排序,你上周找到的匹配结果可能已经消失。找到匹配结果后,立即截图保存结果页面并存档匹配URL。
- 裁剪可能会误导:裁剪草地或天空会返回大量看似可信但毫无意义的匹配结果。应裁剪图片中独特的元素,而非仅仅存在的元素。
- 你可能会泄露信息:将客户的图片上传至商业引擎会向该供应商披露图片内容,并可能暴露你的调查意图。对于敏感材料,搜索时裁剪掉识别性细节——或者不进行搜索。详见。
investigate-without-getting-made
Confidence grading
可信度分级
- Confirmed original source — you have a page that pre-dates every other copy found, hosted by a plausible originator (the photographer, the agency, the subject's own account), with an on-page date corroborated by an independent archive capture, and no larger or earlier version exists.
- Probable — earliest copy is consistent across engines and the host is plausible, but the date rests only on the page's own claim, or the archive has no capture from that period.
- Match only — you can show the image appeared at a given URL by a given date. Say exactly that. Do not upgrade it to "the original".
- Unconfirmed — visual similarity without pixel-level verification, or a face-engine hit, or a single engine's result you could not reproduce.
- 确认原始来源:你找到的页面早于所有其他已发现的副本,由可信的原始发布者(摄影师、代理机构、当事人自己的账号)托管,页面显示日期得到独立存档捕获的证实,且不存在更大或更早的版本。
- 可能为原始来源:各引擎返回的最早副本一致,且托管方可信,但日期仅基于页面自身的声明,或存档没有该时期的捕获记录。
- 仅匹配结果:你能证明该图片在指定日期前出现在指定URL上。如实表述即可,不要将其升级为“原始来源”。
- 未确认:仅视觉相似但未进行像素级验证,或人脸引擎的匹配结果,或单个引擎的结果无法重现。
Worked example
实操案例
An account posts a photo captioned "police raid this morning, [city]".
Full frame in Lens: nothing but generic riot-police stock. Yandex: a dozen
similar police photos, none matching. TinEye: no results — which I note as
uninformative, since TinEye's index is small.
Crop to the shoulder patch and re-search in Bing using region select. It reads as
a municipal force from a different country than the caption claims.
Crop the shop sign in the background, OCR it in Lens, and search the business
name as text. Two hits, both a street in that other country. Dead end on the
image search itself — but the text pivot lands it.
Back to Yandex with the storefront crop: a news gallery from three years earlier,
same street, same shop awning, same barrier arrangement. Photographer credited.
Archive shows a capture of that gallery page two days after its stated date,
which corroborates it.
Conclusion: confirmed that this image was published years before the claimed
event, in another country. Recontextualised, not fabricated. Hand the location to
and the photographer credit to .
geolocate-from-pixelsfind-anyone某个账号发布了一张照片,配文“今早警方突袭,[城市]”。
使用Lens搜索整张图片:仅返回通用防暴警察图库素材。Yandex:返回十几张相似的警察照片,但均不匹配。TinEye:无结果——我记录这一情况,因为TinEye的索引规模较小,无结果不具备参考价值。
裁剪肩章部分,使用Bing的区域选择功能重新搜索。结果显示这是另一个国家的地方警察部队,与配文声称的国家不符。
裁剪背景中的店铺招牌,使用Lens进行OCR识别,然后搜索该商家名称。得到两个结果,均指向另一个国家的某条街道。图片搜索本身陷入僵局——但文字线索找到了突破口。
使用Yandex搜索裁剪后的店面图片:找到三年前的一个新闻图库,包含同一条街道、同一店铺遮阳篷、同一障碍物布置。图片标注了摄影师署名。存档显示该图库页面在其声明日期两天后被首次捕获,证实了日期的真实性。
结论:确认该图片在声称的事件发生多年前就已在另一个国家发布。图片被重新语境化,而非伪造。将地点信息交给工具,将摄影师署名交给工具。
geolocate-from-pixelsfind-anyonePivots
后续操作方向
| What you got | Send to |
|---|---|
| Location, street scene, storefront | |
| Named people, photographer credit | |
| Publishing site or agency domain | |
| Match page that is gone or altered | |
| Caption text, business name, filename slug | |
| Suspected manipulation or generation | |
| Posting account, avatar reused elsewhere | |
| Full photo/video geolocation case | |
Engine-by-engine selection detail:
reference/engine-matrix.md.
| 你获得的信息 | 转交至 |
|---|---|
| 地点、街景、店面 | |
| 提及的人物、摄影师署名 | |
| 发布网站或代理机构域名 | |
| 已消失或被修改的匹配页面 | |
| 图片说明文本、商家名称、文件名标识 | |
| 疑似被篡改或生成的图片 | |
| 发布账号、在其他地方复用的头像 | |
| 完整的照片/视频地理定位案例 | |
各引擎选择的详细说明:
reference/engine-matrix.md。