Getting text into a vector column can be a two-part job. The database holds the vectors. Something outside it, such as a scheduled job or a separate service, turns the text into vectors and writes them back, and that outside piece is where the endpoint credential, the retry logic and the error handling live.
SQL Server 2025 can move the endpoint call inside. CREATE EXTERNAL MODEL registers an embedding endpoint as a database object, and AI_GENERATE_EMBEDDINGS calls it from an ordinary T-SQL statement. That also moves the endpoint credential into the database, and the database server starts making the outbound HTTPS calls itself. A job still has to decide which rows need embedding and when, and to run the statements; what moves is where the call to the endpoint is made.
What Moves Into the Database
Whether anything has to be switched on first depends on the platform. On SQL Server 2025 and Azure SQL Managed Instance, the external rest endpoint enabled option is off by default and has to be set to 1 with sp_configure, followed by RECONFIGURE WITH OVERRIDE. On Azure SQL Database and SQL database in Fabric it is already enabled. After that come a database master key, a database scoped credential, and the model object, which names the endpoint LOCATION, the API_FORMAT, the MODEL_TYPE and the MODEL. The accepted API formats are Azure OpenAI, OpenAI, Ollama and ONNX Runtime, and remote endpoints must use HTTPS with TLS.
Once that exists, the embedding step is a statement: an UPDATE that sets a vector column to AI_GENERATE_EMBEDDINGS over a text column in the same table. Access becomes a set of grants. Creating or changing a model needs CREATE EXTERNAL MODEL or ALTER ANY EXTERNAL MODEL, using one needs EXECUTE on that external model, and sys.external_models lists the models that exist. The list of principals who can send company text to the embedding endpoint can be read with a query, where before it was whichever account the outside job ran under.
The Credential Is Named After the URL
A database scoped credential used by an external model does not get a free-form name. Its name must be a valid URL with no query string. The protocol and fully qualified domain name of the called URL must match those in the credential name, and each part of the called URL path must match the corresponding part of the credential's path. The credential therefore has to point to a path at least as general as the request. Microsoft's own example: a credential created for https://northwind.azurewebsite.net/customers cannot be used for https://northwind.azurewebsite.net.
This matters most with deployment-specific paths. An Azure OpenAI model LOCATION includes the deployment name in its path, so a credential named down to one deployment will not match a model that points at a different deployment on the same host. The documented examples name the credential at the host root, which matches every path below it. Putting the credential name and the model LOCATION side by side before running the CREATE catches the mismatch before it shows up as a failed call.
Retries are configured in two places. retry_count is a whole number from 0 to 10, set on the model inside the PARAMETERS JSON under sql_rest_options. A query can pass its own retry_count in the function's PARAMETERS, and the AI_GENERATE_EMBEDDINGS reference says the query value overrides the one on the model. Two statements that look the same can retry a slow endpoint a different number of times, and nothing on the face of the model definition shows which value a given query used.
Batches, Rollback and Where Failures Show Up
In autocommit mode an UPDATE is its own transaction, and it stays open while every endpoint call and retry for the rows it touches is in progress. If that statement fails and is rolled back, its changes are rolled back with it and the column is left as it was. The database rollback does not undo the endpoint calls already made, so the requests have still been sent and processed by the provider.
Splitting the work into separately committed batches changes that. Batches that committed before a failure stay committed, and the next run can pick up rows whose vector column is still NULL. That makes the batch the unit to plan: how many rows per batch, at what hour, with which retry_count, and how the job records which batch failed. The transaction documentation covers the autocommit and explicit-transaction rules. Inside an explicit transaction, a statement that completes is still not committed until the transaction is.
The job that runs the batches should record whatever the statement returns or raises. For more detail on the database side, AI_GENERATE_EMBEDDINGS has an extended event, ai_generate_embeddings_summary, carrying the REST status code, the errors encountered and the model name used, and external_rest_endpoint_summary adds more request and response detail. A throttling response from the endpoint can be recognized there by its status code, so the session is worth defining before the first large run.
Fixed-Size Chunks Can Split a Word
AI_GENERATE_CHUNKS handles the splitting, and it is deliberately narrow. CHUNK_TYPE accepts one value, FIXED, and CHUNK_SIZE is a count of characters. In the published example, at a chunk size of 50, one chunk ends with "on the top of stee" and the next begins with "p hills such as we see in old missals". The split falls on the character count, mid-word if that is where the count lands.
OVERLAP is the mitigation. It is a percentage of the chunk size, a whole number from 0 to 50, and it defaults to 0. At the default, a sentence that straddles a boundary is divided between two chunks, and a retrieval query can match neither half well. The function also requires database compatibility level 170 or higher. Below that, the Database Engine cannot find it, so a database restored from an older environment fails with a missing-function error that looks like a typo.
When Is the Database the Wrong Place for This?
Generating embeddings in T-SQL fits when the text already lives in SQL Server and the people who would otherwise support a separate embedding service are the ones who already support the database. It fits less well when new text arrives continuously and never pauses long enough for a batch, or when the source documents sit in a file store the database does not read. The embedding model is the other factor. A similarity search compares the query's vector with the stored ones, so both have to come from a consistent embedding model, and a change of model needs a planned transition in which queries keep using the model that produced the vectors they search. One way to run it is a second vector column filled batch by batch while queries stay on the old column and model, with the query model and the searched column switched together once the new column is complete. If the model is expected to change often, that transition work belongs in the decision about where generation runs.
Related Reading
For where the vectors live once they exist, read SQL Server Can Now Power Semantic Search Without a Separate Vector Database. For the retrieval layer underneath both, read Before You Add AI, Fix the Data Retrieval Layer.