Unstructured → structured
Building the knowledge base itself
Our own models, reading scanned pages
Lebanon’s laws weren’t available as clean data. They existed as public scanned pages and scattered documents. We trained our own models to read those pages: detecting titles, article text and boundaries, separating each article, and converting the Arabic into searchable text. That pipeline produced the ~358,000 structured articles Co-Lawyer answers from. Most knowledge systems assume clean data already exists. When it doesn’t, we build it.