AI Comes to Exadata Sizing

The dramatically increased data volume we use for database system sizing has enabled us to use Artificial Intelligence to further improve the process. We previously relied on just a few data points for each database, but now we collect thousands! Read on to learn how we are integrating AI into the process of sizing Exadata for on-premises and Cloud.

Data Explosion

We injected an explosion of data into the sizing process starting in 2019, removing guesswork from the sizing process by making it much more data-driven. Rather than relying on “rule of thumb” approaches, we let the data show us the way forward. We now collect thousands of datapoints for each database and correlate those datapoints across databases to enable virtually ANY consolidation scheme users can imagine.

Oracle Cloud Capacity Analytics

We developed a tool called Oracle Cloud Capacity Analytics (OCCA) that is used internally to size systems for running Oracle Databases. The majority of the databases we size are moving to Oracle Cloud or one of the multi-cloud targets we support, but we also see a stead-stream of databases moving to Exadata on-premises. The OCCA tool gives estimates for the amount of system capacity required to run a set of databases. OCCA relies on data analytics to determine system size required to run those databases now and projecting into the future. The entire process is even more analytical and mathematical than ever before.

We use data science and analytics to size systems that will handle your databases in both Cloud and on-premises. We have also implemented Artificial Intelligence in strategic ways to make the process more effective than ever before.

AI Context Engineering

Our sizing tool begins with careful AI context engineering to ensure the AI is providing accurate answers to questions our users ask, but it also ensures the AI only provides answers in the area of expertise it’s trained on. If you ask the AI questions about things it doesn’t know about, the AI will tell you. For example, here’s what happens when a user asks OCCA why the sky is blue…

Carefully engineering the context the AI works within allows us to focus on training the AI for specific areas of expertise, increasing the quality of answers it gives to users. Context engineering also ensures the AI won’t even try to answer questions it knows nothing about, so we don’t have to worry about it telling users how to rule the world or otherwise do things the AI simply doesn’t handle. The AI we use has great information about sizing such as the definition of a cohort…

The AI we’re using (called “Ask OCCA”) is trained on the OCCA User’s Guide and related content to ensure it gives quality answers that are relevant to our users. It’s interesting that the AI will explain concepts from a different perspective than our technical writers might take on the same topic. Explaining the same topic from a different angle can help users understand the topic even better than before.

AI for Automating Cohorts

Databases are grouped together for consolidation, which saves system resources and reduces labor costs for managing databases. We call these groups of databases “cohorts” because there wasn’t a universally accepted term for this. In short, cohorts are groups of databases that can be consolidated together. The opposite of consolidation is isolation, and it’s important to consider these two ends of the spectrum. Cohorts capture the following:

  • Location of databases
  • Administrative Isolation
  • Security Isolation
  • Blast Radius

Most customers have databases in at least 2 locations such as production and disaster recovery locations. Many cloud vendors also have regions and availability domains for an additional layer of isolation by location, and that scheme is becoming more common in on-premises data centers. Administrative isolation is nearly ubiquitous because most organizations prefer to isolate production from development and test systems. Financial services, banking, and government customers often have specially isolated systems for security purposes, sometimes including multiple separate (and secure) networks. It’s the digital equivalent of security gate, security desk, and locks at each floor of a high rise building. It’s also common for customers to isolate databases from each other to reduce the impact of a failure or what is known as the “blast radius” of a failure.

There are sometimes markers in the data that we collect that suggests what cohort assignments should be. Customers use a wide variety of schemes for organizing their databases including some of the attributes shown below:

  • Database naming conventions
  • Server naming conventions
  • Cluster names
  • Exadata Database Machines used for consolidation
  • Enterprise Manager attributes (cost center, department, lifecycle status, line of business, location)

Artificial Intelligence is an ideal technology to determine what scheme a customer might be using, making cohort assignments much easier to divine from the data. The relationships between Primary and Standby databases can also give us clues about what the cohorts should be, or at least which databases can reside in each cohort. We are looking to AI to help us understand a population of databases and how they should be grouped together before we ask customers how they want to organize their databases. The opposite approach is to simply isolate everything onto dedicated Virtual Machines, but the cost and operational advantages of database consolation are too great to ignore.

AI-Driven Sizing Report Comparisons

Each target platform (Exadata on-premises, Exadata Cloud@Customer, Autonomous Database, etc.) has multiple options, versions, and other settings that affect sizing output. Thousands of data points are ingested for each sizing, and each platform has different characteristics that affect the capacity required to handle the workload. These factors make report comparison extremely complex, so we turned to Artificial Intelligence to make these comparisons useful. AI is able to evaluate the various factors that make one report different from another, allowing us to build a comparison tool that users find helpful.

AI-Driven Gap Handling (AI Code Assist)

Handling gaps in data is a great example of how AI Code Assist has helped us make our data-driven sizing even more effective. Gaps in data occur simply because Oracle Enterprise Manager (OEM) deployments are operational systems used. OEM is used for keeping databases running and helping administrators care for those databases on a day-to-day or even minute-by-minute basis. OEM captures performance metrics for the purpose of viewing and addressing performance problems. However, we use that information for a dramatically different purpose, which is capacity planning and sizing. We often find gaps in the data or “missing metrics” because some databases might be down, might be undergoing an upgrade, or whatever operational state the database could be in when the metrics are extracted.

These missing metrics or data gaps occur in the most recent, most fine-grained data, but we can sometimes use other metrics from prior periods to fill those gaps. The code required to fill these gaps is extremely complex, examining over 10’s or 100’s of thousands of rows with many attributes per row and more than a year of historical metrics for each database. We used AI Code Assist technology to develop the algorithms needed to fill these gaps in the collected metrics. This allowed our relatively small team of developers to deploy the necessary code that is used by hundreds of engineers to build sizing estimates for thousands of customers.

Conclusion

This article should give you a good idea of how we are putting AI to work in the area of Exadata sizing. Whether you are moving to Cloud or deploying Exadata on-premises, AI is helping us do an even more effective job of sizing systems to meet your needs using the inputs and guidance you provide.

Leave a comment