dabla opened a new pull request, #74290:
URL: https://github.com/apache/airflow/pull/74290

   The Microsoft Azure provider has no integration with [Azure AI 
Search](https://learn.microsoft.com/en-us/azure/search/), the full-text and 
vector store behind most retrieval-augmented generation on Azure. A Dag that 
keeps such an index up to date has to build the SDK client from the connection 
itself and repeat, for every write, what the SDK leaves to the caller.
   
   This PR adds `AzureAISearchHook` 
(`airflow.providers.microsoft.azure.hooks.ai_search`) with the 
`azure_ai_search` connection type, on top of the official 
`azure-search-documents` SDK.
   
   ```python
   @task
   async def refresh_index(documents: list[dict]) -> int:
       hook = AzureAISearchHook(index="reports")
       known = {doc["id"] async for doc in hook.search(select=["id"], 
filter="category eq 'report'")}
       return await hook.merge_or_upload(doc for doc in documents if doc["id"] 
not in known)
   ```
   
   **What the hook does**
   
   - Binds an async `SearchClient` to one index, given to the hook or set as 
the default index of the connection. `get_async_conn()` is an async context 
manager that closes the client and its credential.
   - `search()` is an async generator over every matching document (the SDK 
follows the continuation of the service), without the `@search.*` metadata. 
`count()` and `get_document()` (`None` when the key is unknown) cover the other 
reads.
   - `upload()`, `merge()`, `merge_or_upload()` and `delete()` accept any 
iterable and write it in batches of `batch_size` documents, 1000 by default, 
which is the limit of the service per request.
   - A write raises `AzureAISearchIndexingError` naming the rejected keys and 
their reasons. The SDK returns the per-document results of a partially failed 
write (HTTP 207) without raising, so a caller that does not inspect them loses 
documents silently.
   
   **Authentication** follows the other Azure hooks: a service principal 
(client id, secret and `tenantId`), an API key (password without a client id), 
or `DefaultAzureCredential` with the optional `managed_identity_client_id` and 
`workload_identity_tenant_id`. A client id without its secret or tenant is 
refused instead of falling back to another credential.
   
   **Design choices worth a look**
   
   - The hook is async only, like `KiotaRequestAdapterHook`: its use case is an 
async task or a trigger that reads and writes an index next to its other 
awaits. A synchronous twin can follow if there is demand; I left it out to keep 
one code path.
   - The connection extra is read with `get_async_extra_dejson` from 
`common.compat` (#74147), so nothing blocks on the Task SDK from the event 
loop. The provider already carries the `# use next version` marker on 
`common-compat` for it.
   - `azure-search-documents>=11.6.0` is a new dependency of the provider. The 
hook only uses the stable document operations; I ran it against 11.6.0 and 
12.0.0.
   - There is no operator in this PR. The hook is useful by itself from 
`@task`, and operators can be discussed once the hook has settled.
   
   **Tests**
   
   New unit tests cover the endpoint and index resolution, the three 
credentials and the incomplete service principal, search with select, filter 
and order, both count paths, `get_document`, batching for the four write 
methods (also from a generator), the rejected-document error and the validation 
of `batch_size` from the hook and from the connection.
   
   Run locally:
   
   - `pytest 
providers/microsoft/azure/tests/unit/microsoft/azure/hooks/test_ai_search.py`: 
40 passed, 1 skipped (the connection form widget test, because Flask-AppBuilder 
could not be installed on my machine; CI runs it).
   - `mypy` on the changed files, and `prek` for the `pre-commit` and `manual` 
stages from `upstream/main`: all passed.
   - Outside the test suite, the hook with the real SDK client (11.6.0 and 
12.0.0) against a local HTTP fake of the REST API: paged search over 2500 
documents, both counts, `get_document`, the four writes in batches of 1000, and 
a 207 response turned into `AzureAISearchIndexingError`. Not run against a live 
search service from this branch; the hook is ported from a plugin that runs 
against one with `azure-search-documents` 12.0.0.
   
   ---
   
   ##### Was generative AI tooling used to co-author this PR?
   
   - [X] Yes — Claude Code (Fable 5.1)
   
   Generated-by: Claude Code (Fable 5.1) following [the 
guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#gen-ai-assisted-contributions)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to