Skip to content

Text recognition #510

Description

@rossisbudda

Problem/Use Case

I would love to get the true shot numbers from a video burnin, via OCR text recognition or similar. It's a shame to recognise only the cut points, and throw away the text notations within the image.

Solutions

Visually specify a crop region to scan, then make a record of the text within the box per clip and record in some extra attributes.

Proposed Implementation:

Provide a recognise-text flag to enable OCR on the process.

Alternatives:

Run a second process separately to do the OCR

Activity

  1. Breakthrough commented on May 19, 2025

    @Breakthrough
    Owner

    If the intention is to get a time base for the cuts, then doesn't the OCR only need to be done on the first frame for a video? Is there anything preventing automating this in another way, say having a script query the timecode from the first frame, and then shifting the output of PySceneDetect accordingly?

  2. rossisbudda commented on May 20, 2025

    @rossisbudda
    Author

    Yes that's probably a good way to handle it on the first frame of each cut. It is certainly automatable as a separate process. My suggestion is to consider inbuilding this functionality, as I guess most people using PySceneDetect would love it if the cuts picked up useful naming rather than only an indexed identifier (1,2,3,4,5 could also gather sc01_sh010, sc01_sh020, sc02_sh020, etc)

  3. Repository owner locked and limited conversation to collaborators on Mar 2, 2026
  4. converted this issue into a discussion #536 on Mar 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions