Anderson의 관점

선도적인 언어 모델도 ‘해석적’ 전사에 취약하다

mm
Unite.AI를 Google의 선호 소스에 추가
AI-generated image (GPT-2): a construction worker wearing a tool belt stands beneath scaffolding, holding a printed job sheet and scratching his head while facing a large stone wall engraved with a partially completed version of the United States Declaration of Independence in which numerous words have been mistakenly replaced with incorrect alternatives, ending at the word 'foam'. A workbench holding chisels and a hammer stands in the foreground, with stone dust scattered below the unfinished carving. The photorealistic scene is surrounded by a yellow-and-black warning-style border labeled 'AI FANTASY INSIDE' along the top and right edges and 'REALITY OUTSIDE' along the bottom and left edges, with black directional arrows pointing inward.

아마존의 새로운 연구에 따르면 모델이 더智能할수록 Optical Character Recognition을 수행할 때 더 창의적으로 행동할 가능성이 높아진다.

 

Optical Character Recognition (OCR)은 대체로 해결된 문제로 간주된다. 텍스트를 스캔하거나 OCR 시스템에 사진을 제공하면, 시스템은 각 문자를 세세하게 구분하여 가장 좋은 추측을 한다.

Top left, the zoning of characteristic loci in a segmented identified letter (source - https://home.nr.no/~eikvil/OCR.pdf); main image: segmentation result for an institutional repository (Source - https://ijarse.com/ADMIN/admin/postimages/images/fullpdf/1473863345_515_IJARSE.pdf)

Top left, the zoning of characteristic loci in a segmented identified letter (source); main image: segmentation result for an institutional repository (Source).

OCR은 문자를 개별적으로 해석한다. 이러한 문자는 충분히 근접하여 단어 또는 다른 언어적 구조를 구성할 수 있다. 그러나 OCR은 이러한 문제에 관심이 없다. 이러한 문제가 발생할 때, OCR은 이미 다음 작업으로 이동하여 SPELL-체킹은 OCR의 범위 내에 있지 않다.

가장 인기 있는 OCR 시스템 중 하나는 오래된 Tesseract이다. 이는 엄격한 알고리즘 인터프리터로, 다양한 오픈 소스 애플리케이션과 워크플로우에서 널리 사용된다. 모든 운영 체제에서 사용할 수 있는 다양한 ‘플랫’ 알고리즘 OCR 애드온이 존재하며, 이미지 텍스트를 편집 가능한 텍스트로 변환하는 기능은 이제 무료商品의 지위를 달성했다. 그러나 어떤 시스템도 결점이나 특이점이 없다.

추가로, 관련 업무에서 경쟁하는 조직과 상업 제품은 기존 시스템을 AI 지원 문서 관리 시스템으로 적응시켰다. 그러나 핵심 기술은 추정된 그룹의 분할되고 식별된 문자를 내보내는 것이며, 이는 변하지 않는다.

Word Up

OCR은 점점 더 많은 작업과 함께 처리되고 있다. 새로운 세대의 Vision Language Models (VLMs)에 의해 처리된다. 이러한 시스템은 비편집 가능한 텍스트나 래스터화된 텍스트와 대면할 때, 이러한 최신 시스템의 에이전트 버전은 필요한 경우 능숙하게 읽고 음역할 수 있다.

그러나 아마존의 새로운 연구에 따르면, AI가 OCR과 동일한 기능을 복제할 것이라고 생각하는 사람들은 놀랄 수 있다. 새로운 연구는 선호하는 LLM이 전사에서 더 나아가고, 원하는 경우에도 그렇지 않은 경우에도 그렇다. 이러한 행동은 법률, 의료 및 역사적 전사와 같은 분야에서 바람직하지 않은 결과를 초래할 수 있다.

연구자들은 다음과 같이 주장한다:

‘Vision Language Models (VLMs) are increasingly used in place of traditional OCR pipelines for document understanding.

‘[We] show they do not always act as faithful transcribers: when text is imperfect, they often tend to rewrite it into a more plausible form – a behavior that clean-text OCR benchmarks cannot detect.’

연구자들은 새로운 커리된 데이터셋과 벤치마크를 소개한다. 이 데이터셋은 FaithC4로, 15개의 FOSS와 클로즈드 소스 모델을 테스트했다. 여기에는 ChatGPT와 Google Gemini의 버전이 포함된다.

그들은 일반적인 목적의 AI 모델이 전통적인 OCR 시스템보다 불완전한 단어를 더 가능성 있는 단어로 대체할 가능성이 훨씬 더 높다는 것을 발견했다. 때때로, 65%의 뒤섞인 단어가 다시 작성되었으며, 작은 지역 오류도 정확한 텍스트의 다른 부분에서 오류를 유발했다.

짧은 단어는 특히 취약했다. 명시적인 지시가 있지 않으면 수정을 하지 않도록 했지만, 그것은 문제를 완전히 없애지 못했다.

연구는 다음과 같이 말한다:

‘While the rewriting behavior can be beneficial for recovering some genuine errors, it may be problematic for some domains that may require literal transcription, such as legal documents, medical records, and historical manuscripts.

‘Despite its importance, this rewriting behavior has been largely overlooked in VLM research. Existing OCR benchmarks focus on clean text recognition and are therefore insufficient for investigating how models can handle imperfect inputs.’

중심 메시지는 더 발전된 AI 시스템이 알고리즘 OCR 플랫폼이 쉽게 제공할 수 있는 ‘원시’ 오류를 수정하지 않으려는 경향이 있다는 것이다. 많은 경우에, 인간의 해석이 문제를 해결하는 데 필수적이다. 그러나 잘못된 ‘수정’이 이미 적용되었다면, 문제는 대부분의 체크 시스템의 레이더 아래로 미끄러질 것이다.

새로운 연구Do VLMs Read or Rewrite? On Transcription Faithfulness in Vision-Language Models로, 아마존의 5명의 연구자에 의해 수행되었다.

Data and Method

벤치마크 데이터셋을 채우기 위해, 연구자들은 Colossal Cleaned Crawled Corpus를 사용했다. 이는 C4로, 2020년의 웹 크롤링 컬렉션이다. 이 코퍼스에서, 1,455개의 단일 페이지 문서가 선택되었으며, 각 문서는 단일 페이지에 맞게 크기가 조정되었다.

최종 벤치마크에는 500개의 영어 문서, 500개의 중국어 문서 및 455개의 한국어 문서가 포함되어 있으며, 이는 AI 시스템이忠實하게 복사하는지 또는 어떤 방식으로 다시 작성하는지 측정하기 위한 다국어 테스트 베드를 제공한다.

제어된 오류를 도입하여 문서를 테스트하기 전에, 각 문서는 이미지로 렌더링되었다. 약 8%의 적합한 단어가 변경되었으며, 각 단어는 적어도 4개의 문자를 포함했다. 변경되지 않은 버전의 문서는 기준으로 사용되었다.

세 가지 perturbation 변형이 적용되었다: scramble은 단어 내부의 문자를 섞었지만, 첫 번째와 마지막 문자는 유지했다. random은 각 문자를 동일한 문자 집합에서 무작위로 대체했다. visual은 시각적으로 유사한 문자로 대체했지만, 기본 텍스트는 변경했다.

중국어 버전은 scramble 조건을 생략했으며, 이는 중국어의 문자가 공백으로 구분되지 않기 때문이다. 대신, 다른 perturbation에 대해 4개의 문자 창을 사용했다. 한국어의 평균 단어 길이가 더 짧았기 때문에, 수정할 수 있는 단어가 더 적었다.

각 문서는 PDF 페이지로 렌더링되었으며, 경쟁하는 전사 시스템에 전달되었다.

Tests

테스트는 클로즈드 소스 일반 목적 모델을 포함했다. 여기에는 GPT-5.4-mini; Gemini-3-Flash; 및 Gemini-2.5-Flash가 포함되었다. 또한 오픈 소스 모델도 포함되었다. 여기에는 일반 목적 모델 Qwen3.5-4B; Qwen3-VL-4B; InternVL3.5-4B; Gemma4-E4B; 및 Gemma4-E2B가 포함되었다.

또한 OCR 전문 모델이 포함되었다. 여기에는 olmOCR-2-7B; DeepSeek-OCR-2; MinerU2.5-Pro; PaddleOCR-VL-1.5; 및 LightOnOCR-2-1B가 포함되었다. 마지막으로, 전통적인 OCR 엔진 Tesseract와 docTR도 테스트되었다.

세 가지 그룹은 언어에 다른 정도로 의존한다. 일반 목적 AI 모델은 광범위한 언어 지식을 사용한다. OCR 전문 모델은 문서 전사에 대한 작업별 훈련을 강조한다. 전통적인 OCR 시스템은 문자 인식에 의존한다.

두 가지 지표가 채택되었다. 첫 번째는 Word Error Rate (WER)로, 이는 OCR의 표준 지표이다. WER는 시스템이 잘못 전사한 단어의 수를 결정한다. 두 번째 지표는 Edit Distance Similarity (EDS)로, 이는 예측된 텍스트와 원본 텍스트를 문자 수준에서 비교한다.

중국어의 경우, 단어가 공백으로 구분되지 않기 때문에, 문자 수준에서 Character Error Rate (CER)가 사용되었다.

그림 아래에 표시된 바와 같이, 전통적인 OCR 시스템은 모든 세 가지 perturbation 유형에 대해 가장 저항성이 높았다. 이는 이러한 시스템이 주로 문자를 인식하기 때문이다.

Changes in Word Error Rate (left) and Edit Distance Similarity (right) across English, Chinese and Korean after applying the three text perturbation methods. Traditional OCR systems showed the smallest performance degradation overall, OCR-specialized models occupied the middle ground, while general-purpose AI models were typically the most affected by corrupted text.

Changes in Word Error Rate (left) and Edit Distance Similarity (right) across English, Chinese and Korean after applying the three text perturbation methods. Traditional OCR systems showed the smallest performance degradation overall, OCR-specialized models occupied the middle ground, while general-purpose AI models were typically the most affected by corrupted text. Source

OCR 전문 AI 모델은 중간 지대를 차지했다. MinerU와 LightOn은 상대적으로 강력했다. olmOCR과 DeepSeek-OCR는 훨씬 더 민감했다. 이는 문서 전사에 대한 작업별 훈련이 문서 전사를 강화했지만, 전사에서 다시 작성하는 경향을 완전히 제거하지는 못했다.

일반 목적 AI 모델은 전반적으로 가장 나쁨을 보였다. 이러한 모델은 perturbed 텍스트와 대면했을 때, 전사 오류가 크게 증가했다. 대부분의 모델은 전통적인 OCR 시스템이나 OCR 전문 모델보다 훨씬 더 큰 성능 저하를 보였다. 그러나 Gemini-3-Flash는 예외적으로, 전통적인 OCR 시스템과 유사하게 수행했다.

전체적으로, 순위는 각 모델이 언어 지식을 얼마나 의존하는지와 밀접하게 일치했다. 언어 지식에 대한 강한 가정은 다시 작성하는 경향이 더 컸다.

오류 유형의 세부 사항은 다음 표에 표시된 바와 같이, 전통적인 OCR 시스템은 주로 전통적인 전사 오류를 생성했다.

Results from the English 'Scramble' perturbation test, showing how each transcription system distributed its errors across exact rewrites, near-miss transcriptions ('Close'), more substantial transcription errors ('Moderate') and outright incorrect outputs ('Wrong'). Models are ordered by their rate of rewriting corrupted words.

Results from the English ‘Scramble’ perturbation test, showing how each transcription system distributed its errors across exact rewrites, near-miss transcriptions (‘Close’), more substantial transcription errors (‘Moderate’) and outright incorrect outputs (‘Wrong’). Models are ordered by their rate of rewriting corrupted words.

반면, 일반 목적 AI 모델은 훨씬 더 가능성 있는 대체 단어로 오류가 있는 단어를 대체할 가능성이 높았다. 이는 언어적 맥락을 사용하여 다시 작성하고 있음을 강하게 시사한다.

연구는 또한 전사 오류가 오류가 있는 단어에만 국한되지 않는다는 것을 발견했다.

Results from the error propagation analysis, showing per-word error rates on perturbed and unperturbed words after each perturbation type. Altering only around 4.7% of words caused transcription errors to spread into surrounding, unmodified text across many models.

Results from the error propagation analysis, showing per-word error rates on perturbed and unperturbed words after each perturbation type. Altering only around 4.7% of words caused transcription errors to spread into surrounding, unmodified text across many models.

그림 아래에 표시된 바와 같이, 약 4.7%의 단어만 변경해도, 영어 문서에서 오류가 없는 텍스트의 오류율이 크게 증가했다.

Further results from the error propagation analysis, showing the increase in transcription errors on both perturbed and unperturbed words after each perturbation type. General-purpose AI models exhibited the greatest error amplification, with mistakes frequently spreading beyond the originally corrupted words.

Further results from the error propagation analysis, showing the increase in transcription errors on both perturbed and unperturbed words after each perturbation type. General-purpose AI models exhibited the greatest error amplification, with mistakes frequently spreading beyond the originally corrupted words.

일반 목적 VLM은 가장 큰 오류를 보였다. 이러한 모델은 오류가 없는 단어의 오류율이 10배 이상 증가했다. 반면, 전통적인 OCR 시스템은 오류율이 약간만 증가했다.

실제로, 이는 작은 양의 손상된 텍스트가 정확성에 영향을 미칠 수 있으며, 깨끗한 OCR 벤치마크 점수는 실제 전사 정확성을 신뢰할 수 없게 만든다.

마지막으로, 연구자들은 단어 길이가 다시 작성하는 행동에 영향을 미치는지 조사했다. 짧은 단어가 더 취약하다는 것을 발견했다.

짧은 단어는 시각적인 단서가 적기 때문에, VLM은 언어적 맥락에 더 많이 의존한다. 이는 손상된 텍스트가 문자 그대로 전사되는 대신, 가능성 있는 대체 단어로 대체될 가능성이 더 높다.

연구자들은 또한 명시적인 지시가 다시 작성하는 경향을 줄이지만, 완전히 제거하지는 못한다는 것을 발견했다. 이는 다시 작성하는 행동이 현재 VLM의 내재된 특성이라는 것을 시사한다.

연구자들은 다음과 같이 결론을 내린다:

‘For applications that require literal transcription, such as legal document processing, historical manuscript digitization, and forensics, traditional OCR systems or carefully selected OCR-specialized VLMs remain the safest choice, while general-purpose VLMs introduce systematic rewriting that clean-text benchmarks do not capture.’

Conclusion

이 연구는 LLM 문헌에서 반복되는 주제를 다시 강조한다. 즉, 훈련된 VLM 모델은 매우 강력하지만, 중요한 지시에 저항하는 경향이 있다. 이 경우, 이미지에서 발견한 문자를 읽고 해석하라는 지시이다.

더 발전된 모델은 이것을 할 수 없다. 이는 ‘쓸모없는 작업’으로 간주하거나, ‘반쯤 끝난 작업’을 완료하도록 비이성적으로 강제받는 것 같다.

 

* 저자의 인라인 인용을 하이퍼링크로 변환한 것입니다.

2026년 7월 27일 처음 게시되었습니다

기계 학습 분야 작가, 인간 이미지 합성 분야 전문가입니다. Metaphysic.ai의 연구 콘텐츠 책임자였으며, DNEG의 Brahma.ai로의 합병으로 해체될 때까지 그 역할을 수행했습니다.
포트폴리오 사이트: martinanderson.ai
연락처: martin@martinanderson.ai