Files
archivdms/internal/ocr/coords.go
T
patrick 9a24ea29e1 FDN-01: repository & projektgerüst
Git-Repository für bestehenden archivdms-Code initialisiert, Branch-/Commit-Konvention (feature/<ticket>-<slug>-Branches, Ticket-Prefix in Commit-Nachricht) etabliert.
2026-08-11 21:27:53 +02:00

360 lines
15 KiB
Go

// Word-level bounding boxes for OCR text-highlight/overlay (Phase 1 —
// datengrundlage only, see project memory
// project_ocr_textmarkierung_overlay.md). This file adds:
//
// - WordBox / TSV extraction (tesseract's `tsv` output mode)
// - a small geometry-transform mechanism to map word boxes from the
// coordinate space of the final, fully-preprocessed image tesseract
// actually recognized text on, back into the coordinate space of the
// file the frontend actually displays to the user.
//
// Koordinatenraum (why this file exists at all): runTesseract's
// preprocessing pipeline (clampImageSize -> deskewImage -> normalizeContrast
// -> rotateForOSD, see ocr.go) can resize and rotate the image before
// tesseract ever sees it. The frontend, however, always renders the
// untouched original upload (internal/api/document_handlers.go
// handleGetDocumentFile serves doc.StoragePath byte-for-byte; verified
// 2026-07-30 — no transformed copy is ever persisted or served). Word boxes
// from tesseract are therefore in the WRONG coordinate space for direct use
// against the displayed image unless mapped back.
//
// Full forward order in runTesseract (each step optional):
//
// clampImageSize -> deskewImage | deskewImageHough -> normalizeContrast
// -> rotateForOSD -> binarizeImage
//
// and, for image uploads only, one final forward step applied in ocrImage
// AFTER the inversion above: the file's EXIF Orientation (see exif.go), which
// moves the boxes from raw-pixel space into the space the browser actually
// renders. Inversion happens strictly last-forward-step-first
// (mapWordsToOriginal iterates the slice backwards), so any combination —
// e.g. clamp + hough-deskew + OSD 90 degrees + EXIF 6 — composes correctly:
// the geometric chain is undone in reverse, then EXIF is applied once on top.
//
// What is handled exactly vs. approximately:
// - clampImageSize: pure uniform scale -> inverted exactly (simple ratio).
// - normalizeContrast / binarizeImage: no geometry change; still measured
// (geomChain.recordScale) rather than assumed, and dropped as identity.
// - rotateForOSD: our own rotate90CW, always an exact multiple of 90
// degrees -> inverted with pixel-exact integer math (mirrors the forward
// loop in rotateImageFile step for step, no trig/rounding involved).
// - deskewImage (ImageMagick `-deskew`, arbitrary small angle + canvas
// resize to bound the rotated image): inverted via the standard
// rotate-about-center formula using the angle ImageMagick reports via
// `-print "%[deskew:angle]"` plus before/after pixel dimensions. This is
// geometrically the correct construction for a generic "rotate and
// expand canvas" operation. Sign convention reviewed 2026-07-30 against
// ImageMagick's source behaviour: DeskewImage derives the `deskew:angle`
// artifact from the same `degrees` it feeds into the affine matrix
// [[cos,-sin],[sin,cos]], and AffineTransformImage expands the canvas
// symmetrically about the centre (auto-crop off by default) — so the
// centre-to-centre inverse with rad = -angleDeg below is the exact
// transpose. Still not validated against a real deskewed sample's pixel
// output, so treat it as reviewed-but-not-field-verified.
// Per project memory (two prior deskew-angle tuning attempts were tested
// against the doc id 4-9 corpus and rejected — see
// project_deskew_border_trick_tested_negative.md and
// project_deskew_disable_for_photos_tested_negative.md), do NOT blindly
// adjust this formula's sign/rounding by trial and error; instead verify
// against a real deskewed sample (overlay the mapped word boxes on the
// original image) before touching it, and record the result either way.
// - deskewImageHough: the angle is detected by hough_deskew.py but APPLIED
// by `convert -rotate <angle>`, whose sign convention is documented and
// unambiguous (positive = clockwise). The inverse below therefore IS
// verified for this path — the unverified sign caveat above applies only
// to ImageMagick's own `-deskew`/%[deskew:angle] pair.
// - Steps whose geometry cannot be measured (image.DecodeConfig only knows
// the formats this package imports, i.e. JPEG and PNG — a TIFF/BMP/WebP
// upload fails every measurement while ImageMagick still processes it)
// invalidate the whole chain via geomChain, and the document then gets NO
// word boxes. Silently skipping such a step used to leave the remaining
// transforms mapping into a coordinate space that no longer existed.
//
// PDF scope note: for the pdftoppm raster-fallback OCR path, WordBox
// coordinates are mapped back to the *rasterized page PNG's* pixel space
// (post-preprocessing -> pre-preprocessing raster), not further back into
// PDF point/MediaBox coordinate space. The frontend currently renders PDFs
// via the browser's native PDF viewer (iframe over the original file), which
// uses PDF page-coordinate space, not raster pixels — mapping raster pixels
// into that space is a straightforward additional scale step (raster DPI vs.
// MediaBox size, both knowable via pdftoppm's -r 300 and `pdfinfo`) but is
// left for whoever builds the overlay UI in a later phase, since it depends
// on how that phase chooses to render PDF pages (canvas render at a chosen
// DPI vs. native iframe).
package ocr
import (
"image"
"log/slog"
"math"
"os"
)
// WordBox is a single OCR-recognized word with its bounding box, already
// mapped (best-effort — see package doc comment above) into the coordinate
// space of the file the frontend actually displays for the document this
// word was found in.
type WordBox struct {
Text string
Left int
Top int
Width int
Height int
Confidence float64
// Line, Block, Par come straight from tesseract's TSV line_num/block_num/
// par_num columns, useful for later grouping words into lines/paragraphs
// (e.g. for the eventual highlight-overlay UI) without re-deriving that
// from raw positions.
Line int
Block int
Par int
// Page is the 1-based PDF page number this word was found on. Always 1
// for image uploads (a single "page"; there is no page 0 in output).
Page int
}
// geomTransform describes one preprocessing step's effect on image geometry,
// used to invert tesseract's word bounding boxes back towards the originally
// displayed file. See the package doc comment for what is exact vs.
// best-effort here.
type geomTransform struct {
oldW, oldH int
newW, newH int
// angleDeg is the clockwise rotation applied around the image center, in
// degrees. Zero for a pure resize/no-op step.
angleDeg float64
// exact90 marks a rotation known to be an exact multiple of 90 degrees,
// produced by our own rotate90CW (rotateForOSD) — inverted with
// pixel-exact integer math rather than the trig formula used for
// deskew's arbitrary angle.
exact90 bool
}
// invert maps a point (x, y) from the "new" (post-step) image's pixel space
// back into the "old" (pre-step) image's pixel space.
func (t geomTransform) invert(x, y float64) (float64, float64) {
if t.angleDeg == 0 {
if t.newW == 0 || t.newH == 0 {
return x, y
}
scaleX := float64(t.oldW) / float64(t.newW)
scaleY := float64(t.oldH) / float64(t.newH)
return x * scaleX, y * scaleY
}
if t.exact90 {
steps := (int(math.Round(t.angleDeg)) / 90) % 4
if steps < 0 {
steps += 4
}
curW, curH := t.newW, t.newH
cx, cy := x, y
for i := 0; i < steps; i++ {
// Forward step (rotateImageFile/rotate90CW) was, on pixel
// INDICES: src(w,h) -> dst(h,w), src(x,y) -> dst(h-1-y, x).
// mapWordsToOriginal feeds box EDGE coordinates (left..left+width,
// i.e. a continuous [0,w] range, not indices [0,w-1]), so the
// continuous form of the same rotation is used here:
// dst(x,y) = (h - y, x) => src = (cy, curW - cx)
// (Identical convention to applyEXIFOrientation in exif.go; using
// the index form on edge coordinates would shift every box by one
// pixel per rotation step.)
nx := cy
ny := float64(curW) - cx
curW, curH = curH, curW
cx, cy = nx, ny
}
return cx, cy
}
// General case (ImageMagick -deskew): rotation about the image center
// with the canvas expanded to bound the rotated image. Sign convention:
// positive angleDeg == clockwise (ImageMagick `-rotate`), so the inverse
// rotates by -angleDeg about the new centre and re-centres on the old
// canvas. See package doc comment for how far this is verified per path
// (hough: yes; ImageMagick's own -deskew: source-reviewed only).
rad := -t.angleDeg * math.Pi / 180
cxNew, cyNew := float64(t.newW)/2, float64(t.newH)/2
cxOld, cyOld := float64(t.oldW)/2, float64(t.oldH)/2
dx, dy := x-cxNew, y-cyNew
cos, sin := math.Cos(rad), math.Sin(rad)
rx := dx*cos - dy*sin
ry := dx*sin + dy*cos
return rx + cxOld, ry + cyOld
}
// isIdentity reports whether this step changed no geometry at all (same
// dimensions, no rotation) and can therefore be dropped from the chain.
func (t geomTransform) isIdentity() bool {
return t.angleDeg == 0 && t.oldW == t.newW && t.oldH == t.newH
}
// mapWordsToOriginal applies transforms in reverse (last-applied-preprocessing-
// step-first) order, mutating words in place to convert their bounding boxes
// from final-tesseract-image space into the coordinate space of the file
// before any of these transforms ran.
//
// ALL FOUR corners are inverted, not just top-left/bottom-right. That matters
// as soon as a non-90-degree rotation (deskew) is in the chain: under a
// rotation the two opposite corners alone no longer span the rotated
// rectangle's axis-aligned bounding box — for a typical 2-3 degree deskew the
// resulting box is systematically too narrow/short and offset, and at angles
// approaching 45 degrees it collapses towards zero size. The result here is
// the true axis-aligned bounding box of the back-rotated word quad, which is
// what the frontend overlay draws.
func mapWordsToOriginal(words []WordBox, transforms []geomTransform) {
if len(transforms) == 0 {
return
}
for i := range words {
l, t := float64(words[i].Left), float64(words[i].Top)
r, b := float64(words[i].Left+words[i].Width), float64(words[i].Top+words[i].Height)
corners := [4][2]float64{{l, t}, {r, t}, {r, b}, {l, b}}
for c := range corners {
x, y := corners[c][0], corners[c][1]
for j := len(transforms) - 1; j >= 0; j-- {
x, y = transforms[j].invert(x, y)
}
corners[c][0], corners[c][1] = x, y
}
minX, maxX := corners[0][0], corners[0][0]
minY, maxY := corners[0][1], corners[0][1]
for c := 1; c < 4; c++ {
minX = math.Min(minX, corners[c][0])
maxX = math.Max(maxX, corners[c][0])
minY = math.Min(minY, corners[c][1])
maxY = math.Max(maxY, corners[c][1])
}
words[i].Left = int(math.Round(minX))
words[i].Top = int(math.Round(minY))
words[i].Width = int(math.Round(maxX - minX))
words[i].Height = int(math.Round(maxY - minY))
}
}
// geomChain collects the geometry-changing preprocessing steps of a single
// runTesseract pass, so word boxes can be inverted back into the source
// image's coordinate space afterwards.
//
// The important property it enforces (this was a real, silent bug before):
// a preprocessing step that DID change geometry but whose geometry could not
// be measured must invalidate the whole chain, not just be skipped. Skipping
// it leaves the remaining transforms mapping into a coordinate space that no
// longer exists, and the frontend then draws a confidently wrong overlay.
// The realistic trigger is an upload format image.DecodeConfig cannot read:
// this package only registers image/jpeg and image/png, so TIFF/BMP/WebP/GIF
// uploads (all accepted as image/*) fail every decodeImageDims call while
// ImageMagick happily processes them. Rather than misplace boxes we return
// none for those documents.
type geomChain struct {
steps []geomTransform
broken bool
log func(level slog.Level, msg string, args ...any)
}
// recordScale books a step that may only scale the image uniformly
// (clampImageSize) or must not change geometry at all (normalizeContrast,
// binarizeImage). Identity steps are dropped.
func (c *geomChain) recordScale(step, oldPath, newPath string) {
t, ok := buildScaleTransform(oldPath, newPath)
if !ok {
c.fail(step, "image dimensions unreadable (unsupported format for image.DecodeConfig?)")
return
}
if t.isIdentity() {
return
}
c.steps = append(c.steps, t)
}
// recordRotation books a rotation step (deskewImage/deskewImageHough/
// rotateForOSD). A reported angle of 0 combined with changed dimensions means
// the angle was lost (e.g. an ImageMagick build not populating
// %[deskew:angle]) while a rotation really was applied — unrecoverable, so
// the chain is invalidated instead of silently mapping with angle 0.
func (c *geomChain) recordRotation(step, oldPath, newPath string, angleDeg float64, exact90 bool) {
t, ok := buildRotationTransform(oldPath, newPath, angleDeg, exact90)
if !ok {
c.fail(step, "image dimensions unreadable (unsupported format for image.DecodeConfig?)")
return
}
if angleDeg == 0 && (t.oldW != t.newW || t.oldH != t.newH) {
c.fail(step, "rotation applied but angle unknown (0) — cannot invert")
return
}
if t.isIdentity() {
return
}
c.steps = append(c.steps, t)
}
func (c *geomChain) fail(step, reason string) {
c.broken = true
if c.log != nil {
c.log(slog.LevelWarn, "ocr word boxes disabled: preprocessing geometry not invertible",
"step", step, "reason", reason)
}
}
// transforms returns the collected chain; ok is false when any step could not
// be recorded reliably, in which case callers must not emit word boxes at all.
func (c *geomChain) transforms() ([]geomTransform, bool) {
if c.broken {
return nil, false
}
return c.steps, true
}
// decodeImageDims returns the pixel width/height of the image at path
// without decoding full pixel data (image.DecodeConfig only reads the
// header).
func decodeImageDims(path string) (w, h int, err error) {
f, err := os.Open(path)
if err != nil {
return 0, 0, err
}
defer f.Close()
cfg, _, err := image.DecodeConfig(f)
if err != nil {
return 0, 0, err
}
return cfg.Width, cfg.Height, nil
}
// buildScaleTransform records a pure-resize geometry step (clampImageSize)
// by decoding both images' dimensions. ok is false if either image's
// dimensions cannot be read, in which case the caller should skip recording
// a transform (best-effort, same tolerance as the rest of this package).
func buildScaleTransform(oldPath, newPath string) (geomTransform, bool) {
oldW, oldH, err := decodeImageDims(oldPath)
if err != nil {
return geomTransform{}, false
}
newW, newH, err := decodeImageDims(newPath)
if err != nil {
return geomTransform{}, false
}
return geomTransform{oldW: oldW, oldH: oldH, newW: newW, newH: newH}, true
}
// buildRotationTransform records a rotation geometry step (deskewImage or
// rotateForOSD) by decoding both images' dimensions plus the rotation angle
// applied. exact90 distinguishes rotateForOSD's pixel-exact 90-degree
// rotations from deskewImage's arbitrary-angle, best-effort inverse.
func buildRotationTransform(oldPath, newPath string, angleDeg float64, exact90 bool) (geomTransform, bool) {
oldW, oldH, err := decodeImageDims(oldPath)
if err != nil {
return geomTransform{}, false
}
newW, newH, err := decodeImageDims(newPath)
if err != nil {
return geomTransform{}, false
}
return geomTransform{
oldW: oldW, oldH: oldH,
newW: newW, newH: newH,
angleDeg: angleDeg,
exact90: exact90,
}, true
}