Trust

Trust and Security FAQ

Cerenovus handles a great deal of sensitive customer data, so trust and security are among our highest priorities. We’ve compiled the questions we’re most often asked by IT and security teams, all in one place.

1. What data do we collect? — Flexible

We offer to connect to documents, communications, and databases, but the specifics vary considerably from company to company. No two companies have exactly the same data sources they want to connect. There’s always a tradeoff between better results and what you’re comfortable with in terms of data privacy. It’s also important to keep in mind any specific goals you have when choosing which data you’d like us to analyze — we only want the data we need to give you the results you want, no more, no less.

2. Where does the data live and how long is it retained? — Flexible

Typically, we recommend a single-tenant deployment on a cloud machine run by us. This is very secure, given that we can configure the network to only allow your machines (for data upload) and our machines (for installation and telemetry) to connect to the deployment. If further security is necessary, we can explore bring-your-own-cloud or full on-prem deployments.

How long data is retained is entirely up to you. We keep it only for as long as we’re actively using it, and we can work within your existing constraints.

3. Is my data used to train models? — No

We do not use customer data to train our own models, and we have zero-data-retention agreements with OpenAI and Anthropic, so we can guarantee that your data is never used to train ML models.

4. What’s the state of your compliance certifications?

We are in the process of being certified for SOC 2 Type II and ISO 27001 — both of these processes take several months to complete, but we’re moving as fast as we can. In the meantime, we’re happy to sign any NDAs or other agreements as necessary.

5. How does your product work, exactly, from a technical standpoint?

Our product revolves around building context-based mechanisms that allow LLMs to work with massive corpora of information. Without going into proprietary detail, this means a combination of high-quality retrieval mechanisms (we use a hybrid graph + embeddings architecture), and a custom synthesis pipeline that combines exhaustive data taxonomization with a self-improving loop to slowly build up and refine its understanding of the corpus. We also have systems that pull in external information to augment whatever data you give us, allowing us to make use of public information when it is useful. We typically use frontier models running in our own harness, but for special circumstances we can also use local models running on your infrastructure.

In short:

  • We use context mechanisms (not training our own models).
  • We can pull in outside data to supplement.
  • Our deliverable is primarily the finished, verified reports, but depending on the details we could also explore agent integrations.
  • The strongest part of our product (from a technical perspective) is the synthesis component — the ability to take data that’s scattered, organize it, and draw conclusions from it.

If you have any further questions about the technical parts of the product, either book a demo and indicate that you’re interested in talking with our technical staff, or reach out directly to our CTO at ollie@cerenovus.ai.