OpenAI has pushed the GPT family another step forward with GPT-6.1 Sol, a model positioned for complex coding, computer use, and professional work while targeting substantially lower cost than GPT-6 Astra. OpenAI says GPT-6.1 Sol delivers near-Astra performance for these workloads, with standard API pricing of $2 per million input tokens and $10 per million output tokens, while cached input is priced at $0.10 per million tokens. The model has a 1.05 million-token context window and supports up to 128,000 output tokens.
That is important for more than model benchmarking.
It points toward a larger change in how companies should think about AI systems.
As models become more capable, the model itself becomes only one part of the system.
The harder question becomes:
What information can the model access, how quickly can it retrieve it, what version of that information is authoritative, and what is it allowed to see?
That moves the center of gravity toward data infrastructure.
GPT-6.1 Sol is not just about a larger model
The GPT-6 family is broader than a single model.
OpenAI currently describes GPT-6 Astra as its flagship model for demanding work, GPT-6.1 Sol as a lower-cost option aimed at complex work, and GPT-6 Luna as an efficient model for focused, high-volume workloads. OpenAI’s developer guidance frames model selection around the tradeoff between reasoning capability, latency, and cost rather than treating one model as universally appropriate.
GPT-6.1 Sol is particularly interesting because OpenAI is targeting work that moves beyond simple question answering.
The official positioning includes:
- agentic coding
- computer use
- professional work
- tool use
- structured outputs
- web search
- file search
- computer interaction
- hosted execution and other tools through the broader Responses API ecosystem
In practical terms, this changes the architecture of an AI application.
A traditional application might look like:
user → application → database → response
An AI-native application increasingly looks more like:
user → model → reasoning → retrieval → tools → enterprise systems → actions → result
The model may be the reasoning engine, but it still needs the rest of the system to be useful.
That distinction is becoming increasingly important.
The real enterprise AI problem is context
A highly capable model does not automatically know a company’s private information.
It does not automatically know:
- the latest internal pricing policy
- this quarter’s inventory
- a customer’s current contract
- an organization’s private engineering documentation
- internal incident history
- production telemetry
- current financial records
- proprietary research
- internal compliance procedures
- application-specific business rules
Those things live somewhere else.
They live in databases, warehouses, object stores, document repositories, applications, APIs, event streams, file systems, and operational systems.
So the practical question for enterprise AI becomes:
How does the model get the right context from those systems at the right moment?
This is where retrieval enters the architecture.
OpenAI’s own developer documentation describes retrieval as semantic search over a company’s data and explains that its Retrieval API is powered by vector stores. OpenAI’s file-search tooling similarly allows models to search previously uploaded knowledge bases using semantic and keyword search.
The model is therefore not replacing the data system.
It is increasingly working through the data system.
A large context window is not a database
GPT-6.1 Sol has a 1.05 million-token context window. That is significant because models can process much larger amounts of information within a single interaction than earlier systems.
But a context window and a database solve different problems.
A context window answers:
How much information can the model process in this interaction?
A data system answers:
Where does the information live, how is it organized, how is it updated, who can access it, and how can it be found reliably?
Those are different questions.
Imagine a company with:
- 20 years of contracts
- billions of transaction records
- millions of support messages
- product documentation
- engineering logs
- telemetry
- images and files
- customer profiles
- financial records
Putting everything into a model context is neither a sensible operational design nor a substitute for persistent data infrastructure.
The useful pattern is selective access.
Instead of sending everything to the model, the system finds the information relevant to the task.
That is retrieval.
Why vector databases matter to GPT applications
One of the most discussed components of the modern AI stack is the vector database.
The basic idea is straightforward.
Text, documents, images, or other information can be represented as vectors — numerical representations that allow systems to compare semantic similarity.
Instead of asking:
Does this document contain exactly these words?
a semantic retrieval system can ask:
Which pieces of information are most conceptually relevant to this question?
This is particularly useful when users phrase a request differently from the way the source material was written.
For example:
“Why did our enterprise customers stop renewing last quarter?”
The relevant information may not contain the exact phrase “stop renewing.”
It could exist across:
- customer interviews
- account notes
- support conversations
- renewal records
- product usage data
- sales notes
- incident reports
Semantic retrieval helps locate related information.
OpenAI explicitly recommends vector databases for quickly retrieving nearest embedding vectors at scale and provides vector-store-based retrieval and file-search capabilities in its platform.
That is why the growth of capable LLMs and the growth of vector search are closely connected.
The better the model becomes at using context, the more valuable it becomes to retrieve the right context.
But vector search is only one part of retrieval
There is an important misconception here.
Enterprise retrieval is not simply:
text → embedding → nearest vector → answer
Real enterprise systems usually need more.
A production retrieval layer may have to combine:
semantic similarity
with
keyword matching
with
metadata filters
with
permissions
with
freshness
with
structured queries
with
ranking
with
business rules
OpenAI’s current retrieval APIs include filtering and ranking controls for vector-store search, while file search combines semantic and keyword search.
Consider an internal banking application.
A user asks:
“Show me the recent risk events affecting European commercial accounts.”
A useful retrieval system needs more than semantic similarity.
It may need to understand:
- what “recent” means
- which records represent risk events
- which customers belong to the commercial segment
- which accounts are in Europe
- whether the user has access to those records
- whether the information is current
- whether a structured database query should be combined with document retrieval
The problem is therefore broader than vector search.
It is a data access problem.
RAG is becoming a data architecture problem
Retrieval-Augmented Generation, usually called RAG, became popular as a way to give language models access to information beyond their trained knowledge.
The basic pattern is familiar:
Question
↓
Retrieve relevant information
↓
Add context
↓
Generate response
But enterprise systems quickly expose the limits of a simplistic RAG architecture.
A real application may need:
User
↓
Identity + permissions
↓
Query understanding
↓
Structured data retrieval
↓
Semantic retrieval
↓
Document retrieval
↓
Ranking
↓
Context assembly
↓
Model reasoning
↓
Tool execution
↓
Audit + observability
That looks much less like a simple chatbot.
It looks like distributed infrastructure.
This is one of the most important shifts happening around modern AI applications.
RAG is moving from a model feature toward a systems architecture.
The enterprise data stack is getting more complicated
There was a time when a typical application could be described with a relatively small database stack.
Today, an AI application may involve:
- PostgreSQL or another relational database
- document storage
- object storage
- vector databases
- search infrastructure
- data warehouses
- streaming systems
- graph databases
- caching
- feature stores
- ETL pipelines
- event buses
- observability systems
- model APIs
Each tool can be valuable.
The issue is not that these technologies exist.
The issue is what happens between them.
A company might have customer records in one system.
Documents in another.
Embeddings in a vector database.
Events in a streaming platform.
Analytics in a warehouse.
Operational state in a relational database.
The AI application must then understand all of those representations.
That introduces:
- synchronization
- duplication
- ingestion
- indexing
- schema translation
- consistency questions
- access-control boundaries
- operational complexity
The model may become smarter while the data architecture becomes more fragmented.
That is an uncomfortable tradeoff.
Better models increase the value of better data
This is the central idea.
A model can improve reasoning.
It can improve coding.
It can improve computer use.
It can improve tool selection.
It can improve autonomous workflows.
But the model still needs information about the world in which it is operating.
For enterprises, that world is mostly proprietary.
The value may be inside:
customer data
operational data
financial data
engineering data
documents
contracts
events
telemetry
knowledge bases
internal applications
The model provides intelligence.
The data provides context.
And the infrastructure connects the two.
That means progress in AI models can actually make investment in data infrastructure more important, not less.
GPT-6.1 and the rise of AI agents
The significance becomes clearer when AI moves from answering questions to taking actions.
An AI assistant that writes a paragraph needs limited access to company infrastructure.
An AI agent that:
- investigates a customer issue,
- reads the account history,
- checks recent transactions,
- searches internal documentation,
- reviews support tickets,
- identifies the likely cause,
- creates a remediation plan,
- updates another system,
needs considerably more.
It needs access to systems of record.
It needs identity.
It needs permissions.
It needs retrieval.
It needs tool execution.
It needs observability.
It needs durable state.
It needs the ability to distinguish current information from stale information.
This is why the infrastructure around AI agents is becoming as important as the model itself.
OpenAI’s current enterprise platform direction reflects this shift. Its Frontier platform describes connections to systems of record, enterprise data, agent execution, identity, permissions, observability, and governance alongside model intelligence.
The architecture is increasingly:
Model + Data + Tools + Policy + State
rather than simply:
Model
Enterprise AI needs more than unstructured documents
A lot of early RAG systems focused almost entirely on documents.
That made sense.
PDFs, manuals, knowledge bases, wiki pages, support articles, and internal documentation are easy to imagine as a retrieval corpus.
But enterprise applications rarely consist only of documents.
Consider an industrial company.
A useful AI system may need:
- equipment manuals
- maintenance records
- machine telemetry
- sensor events
- work orders
- parts inventory
- inspection images
- operator reports
- asset relationships
A financial organization may need:
- transactions
- customer profiles
- contracts
- reports
- event histories
- risk records
- regulatory documents
A modern software company may need:
- application state
- logs
- metrics
- traces
- source code
- incident records
- customer conversations
- product analytics
The AI application therefore needs multiple data shapes.
This is where the traditional boundary between “database,” “search system,” “document store,” “vector database,” and “analytics system” becomes increasingly interesting.
The opportunity for unified data infrastructure
There is a larger architectural question underneath all of this:
What if different data models did not have to become entirely separate infrastructure silos?
A unified data infrastructure approach does not mean every workload must use the same index or the same query algorithm.
It means the underlying system can provide a common foundation for different forms of data.
That foundation can potentially include:
- transactions
- structured rows
- JSON documents
- time-series data
- vectors
- graph relationships
- objects and blobs
- distributed state
The important word is foundation.
Different models can have specialized access paths while still sharing deeper infrastructure.
That can mean shared concerns around:
- storage
- durability
- recovery
- metadata
- transactions
- permissions
- lifecycle
- observability
- replication
- deployment
The purpose is not to eliminate specialization.
It is to reduce unnecessary duplication between specialized systems.
Where PLOMID fits
This is the architectural problem PLOMID is exploring.
PLOMID is building a unified data infrastructure platform that brings multiple data models and workloads toward a common underlying data layer.
Today, the platform focuses on:
SQL
JSON
time-oriented data
Vector, graph, and object workloads are part of the development roadmap rather than being presented as completed capabilities.
The underlying idea is deliberately broader than “build another database.”
The question is:
Can modern applications work with different data shapes without turning every boundary between those shapes into another infrastructure project?
That question becomes especially relevant as AI applications combine operational records, documents, semantic search, events, and other forms of enterprise information.
The AI model does not remove that problem.
In many cases, it makes the problem more visible.
A practical architecture for GPT-powered applications
A production enterprise AI application can be thought of as several layers.
1. Source data
The underlying systems of record:
- SQL tables
- documents
- files
- event streams
- telemetry
- business applications
- external data
2. Data access
The systems that make that information searchable and usable:
- SQL queries
- indexes
- vector search
- keyword search
- graph traversal
- filters
- metadata
- aggregations
3. Retrieval
The layer that decides what information is relevant to the current task.
This may involve:
- semantic retrieval
- exact matching
- structured predicates
- time windows
- ranking
- access-control filtering
4. Context assembly
The system selects the information that should actually reach the model.
This is an important step.
More context is not automatically better context.
The goal is:
relevant + current + authorized + compact
5. Model reasoning
GPT-6.1 Sol or another model processes that context and reasons about the task.
6. Tools and actions
The model may then:
- query another system
- call an API
- run code
- inspect a file
- operate software
- trigger a workflow
7. Governance
Everything sits inside policies governing:
- identity
- permissions
- auditability
- privacy
- retention
- residency
- security
- observability
That is an AI infrastructure stack.
Not just an AI model.
Security becomes more important as models become more capable
There is another consequence of more capable AI that deserves more attention.
A system that can reason across a larger amount of information is also a system that may be able to make use of information that it should not have access to.
That makes authorization part of the AI architecture.
The system needs to answer:
Who is asking?
What can they access?
What information can the model access on their behalf?
Which tools can the model execute?
What actions require approval?
What should be logged?
OpenAI’s enterprise security documentation says business data is not used to train its models by default and describes encryption, retention controls, access management, audit capabilities, data residency options, and other enterprise controls.
Those controls are important, but they also highlight a broader architecture principle:
AI security starts before the model receives the data.
Data residency is part of AI architecture
For global enterprises, location matters.
A company may have different requirements for:
- where data is stored
- where data is processed
- where models are invoked
- which employees can access it
- which country or region governs the data
- which copies may exist outside a jurisdiction
OpenAI’s GPT-6.1 Sol documentation currently lists US and EU data residency support, with additional eligibility conditions described in its platform documentation.
That is one reason “AI infrastructure” increasingly overlaps with questions traditionally associated with databases and distributed systems.
Data location is not just a deployment detail.
It can become a policy.
The role of a vector database may also evolve
The classic idea of a vector database is:
store embeddings → perform similarity search
That remains useful.
But enterprise systems are likely to demand more context around those vectors.
For example:
vector
document
metadata
permissions
timestamp
customer
transaction
relationships
A vector result by itself may be insufficient.
Imagine retrieving a support document that is semantically perfect but belongs to an obsolete product version.
The retrieval system must understand more than similarity.
It needs context.
This is why the future of AI retrieval may look increasingly hybrid:
semantic retrieval + structured filtering + temporal reasoning + authorization + ranking
The vector is an important signal.
It is not the entire data model.
Why this matters for database architecture
The database industry has spent decades optimizing for:
- transactions
- consistency
- indexes
- query planning
- storage
- durability
- replication
- analytical processing
AI adds another dimension:
context retrieval
The interesting architectural question is whether AI workloads should remain a layer sitting on top of completely separate infrastructure or whether more of those capabilities can eventually live closer to the underlying data layer.
That does not imply that every system should become one giant database.
Different workloads have genuinely different requirements.
But it does suggest that the boundaries between:
database
search
vector database
document store
analytics
graph
and
AI retrieval
may increasingly become architectural questions rather than fixed product categories.
GPT-6.1 raises the value of the data layer
There is a tempting way to interpret every new AI model release:
better model → better AI
That is true, but incomplete.
For enterprise systems, the more useful equation may be:
better model + better context + better retrieval + better data = better application
And as models become more capable:
the value of the context increases.
That is the part that matters.
A company does not usually have a competitive advantage because a language model can answer a generic question.
Its advantage may come from the information the model can use:
- customer history
- operational knowledge
- proprietary research
- internal processes
- specialized datasets
- business rules
- real-time events
The model can reason over that information.
But the information has to be available first.
The next AI infrastructure race may be underneath the model
The first generation of generative AI infrastructure focused heavily on model serving.
Then attention moved toward:
- GPU infrastructure
- inference optimization
- embeddings
- vector databases
- RAG
- agent frameworks
- model orchestration
The next layer is increasingly about the data foundation beneath all of them.
A useful enterprise AI system may eventually need a single coherent way to work with:
structured data
documents
vectors
events
relationships
objects
real-time state
The challenge is not merely storing these things.
It is making them available with useful semantics and predictable guarantees.
What GPT-6.1 means for builders
For application developers, the implication is practical.
Do not design the AI architecture around the model alone.
Design around the complete information flow:
Where is the data?
How fresh is it?
How is it indexed?
How is it retrieved?
How is access controlled?
How is it combined?
How is it updated?
How is it audited?
What happens when the model takes action?
Those questions will often matter more to a production AI system than the choice between two models whose benchmark scores are relatively close.
The model is increasingly becoming one component in a larger system.
The bigger picture
GPT-6.1 Sol is significant because it continues a trend that is bigger than the model itself.
AI is moving from:
generate text
toward:
understand context
then toward:
use tools
and increasingly toward:
complete work
As that happens, enterprise AI becomes deeply dependent on the systems that provide information and authority.
That means the AI stack increasingly looks like:
Models
↓
Reasoning
↓
Retrieval
↓
Data infrastructure
↓
Systems of record
↓
Real-world operations
The model may be the visible part.
The data layer is the foundation underneath it.
Final thought
The most interesting consequence of GPT-6.1 may not be another improvement in model capability.
It may be what that improvement exposes.
The better the intelligence becomes, the more valuable the context becomes.
And context comes from data.
That means the future of AI will not be determined only by who builds the most capable models.
It will also depend on who builds the infrastructure that allows those models to access the right information — with the right performance, consistency, security, permissions, and operational guarantees.
AI provides the intelligence.
Data provides the context.
And increasingly, the infrastructure connecting the two is becoming the real system.
PLOMID is exploring that infrastructure layer: a unified data foundation for modern applications where SQL, JSON, time-oriented data, and future workloads such as vectors and graphs can move toward a common underlying system.
Explore PLOMID
Read the architecture
Explore the roadmap
Read more from PLOMID