This week, news emerged about open-source tools that aim to reduce reliance on Nvidia's CUDA ecosystem, highlighting a broader shift in how AI infrastructure is being built. While the focus was on compute and communication libraries, we found ourselves thinking about the layers beneath - the data substrate that powers AI applications and the tradeoffs between vector databases and data substrates in real-world use cases.
At Apex Grid, we're building a regtech tool for Nigerian microfinance banks, where compliance and traceability are not just features - they're requirements. One of the core challenges we faced was ensuring that every AI inference could be traced back to a specific version of training data, with citable sources and audit logs. This is where a data substrate - a versioned, citable, and queryable reference layer for AI - becomes essential.
Vector databases and RAG (Retrieval-Augmented Generation) systems are powerful for specific use cases, like semantic search or chatbot memory. But they're not designed for the kind of traceability and version control that regulatory environments demand. For instance, a vector DB might store embeddings of documents, but it doesn't track which version of a document was used in a particular inference. RAG-over-files might pull data from a corpus, but it lacks the fine-grained metadata required for compliance audits.
To illustrate, here's a simplified example of how we model data in our regtech tool using a data substrate approach. We maintain a versioned dataset of loan applications, each with a unique identifier, timestamp, and provenance:
class LoanApplication:
def __init__(self, application_id, timestamp, data, version, provenance):
self.application_id = application_id
self.timestamp = timestamp
self.data = data
self.version = version
self.provenance = provenance
def get_version(self):
return self.version
def get_provenance(self):
return self.provenance
Each time a new version of a loan application is submitted, it's stored as a new entry in the dataset. When an AI model is trained or queried, it references the specific version and provenance of the data used. This ensures that any regulatory inquiry can be traced back to the exact data and version used in the inference.
Vector databases, on the other hand, are better suited for use cases where the primary concern is speed and similarity search. For instance, a customer support AI that needs to quickly find similar past interactions would benefit from a vector DB. But in our case, the need for traceability and auditability made a data substrate the right choice.
That said, there are tradeoffs. A data substrate requires more storage and careful design to manage versioning and provenance. It also introduces complexity in querying and indexing, especially when dealing with large datasets. Vector DBs are more performant for certain types of queries but lack the metadata richness of a data substrate.
We're now working on a hybrid approach: using a data substrate for core compliance and audit use cases, while leveraging vector DBs for parts of the system that require high-speed similarity search, like fraud detection. This allows us to maintain the traceability we need without sacrificing performance.
What's next? We're exploring how to integrate this hybrid model with on-device AI, ensuring that even when data is processed locally, it remains citable and traceable. We're also looking into how to make the data substrate more interoperable with existing AI frameworks - a challenge that's only going to become more pressing as the industry moves away from monolithic systems toward modular, auditable components.
Top comments (0)