Hearleaf
A consent-first mobile reader that turns photographed printed pages into progressively generated, locally cached audio in a voice the user controls.
- Context
- Printed pages are not equally accessible, while page images and personal voices demand careful handling.
- Scope
- Mobile capture, OCR review, voice authorization, speech generation, playback, cloud delivery, and operations.
- Role
- Product owner and hands-on engineer across the mobile app, backend, speech pipeline, and infrastructure.
- Constraints
- Consent must precede page upload or voice use, and long-running jobs must survive retries without blocking playback.
- Decisions
- Default to reviewed on-device OCR, authorize voices before durable work, and generate immutable verified audio chunks.
- Outcomes
- Shipped a live Android product and simplified production from ten Cloud Run services to one CPU and one GPU service.
- Lessons
- Trust boundaries and recovery paths belong in the architecture, not in post-launch policy.
Under the hood
A reviewed path from page to playback
- 01Capture
- 02Review
- 03Authorize
- 04Generate
- 05Listen