Open Closed

TenantCreatedEto handler fails intermittently due to distributed lock contention between multiple EfCoreDatabaseMigrationEventHandlerBase instances #10502


User avatar
0
laura created

ABP Framework version: 9.1.1

Template: Microservice (with multi-tenancy)

When creating a new tenant in a microservice template, the admin user is sometimes created and sometimes not. The behavior is non-deterministic.

Root cause analysis:

In the Identity Service, there are ~7 registered IDistributedEventHandler<TenantCreatedEto> implementations — one from our custom IdentityServiceDatabaseMigrationEventHandler (which seeds the admin user) and 6 from ABP's built-in EF Core modules (SaasEntityFrameworkCoreModule, AbpPermissionManagementEntityFrameworkCoreModule, AbpFeatureManagementEntityFrameworkCoreModule, AbpSettingManagementEntityFrameworkCoreModule, AbpIdentityProEntityFrameworkCoreModule, AbpOpenIddictProEntityFrameworkCoreModule).

When a TenantCreatedEto event is received, all 7 handlers are dispatched nearly simultaneously (within milliseconds). Each handler calls MigrateDatabaseSchemaAsync(), which attempts to acquire a distributed lock via Redis. Only one handler can acquire the lock at a time — the remaining 6 time out and throw an exception.

In the custom handler's HandleEventAsync override, both MigrateDatabaseSchemaAsync() and the admin seeding (_dataSeeder.SeedAsync()) are wrapped in a single try/catch. If the distributed lock times out during migration, the entire block fails — including the admin user seeding. The exception is caught by HandleErrorTenantCreatedAsync which does not rethrow, so the inbox processor considers the event "processed" and does not retry.

Evidence from AbpEventInbox table:

Each tenant creation generates 7 inbox records with the same MessageId, all marked as Processed = true. Example for tenant "lbb":

  • 7 records, same MessageId 3a1fe4020958a96fbf4bcf7f97ce85e3
  • CreationTime spread: 08:02:16.189 to 08:02:16.196 (7ms total)
  • ProcessedTime spread: 08:02:16.650 to 08:02:16.857 (207ms total)

Impact: In a shared-database scenario, the admin user is intermittently not created for new tenants, making the tenant unusable.

Expected behavior: The admin user should always be created regardless of distributed lock contention, since for shared-database tenants the schema migration is unnecessary (the schema already exists from the host migration).

Markdown supported.
Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)

9 Answer(s)
  • User Avatar
    0
    maliming created
    Support Team Fullstack Developer

    hi

    I will check this case asap.

    Thanks.

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
  • User Avatar
    0
    maliming created
    Support Team Fullstack Developer

    hi

    The HandleErrorTenantCreatedAsync will try to publish the event again, and your event handler will execute again.

    BTW, it will delay random 5-15s.

    protected virtual async Task HandleErrorTenantCreatedAsync(
        TenantCreatedEto eventData,
        Exception exception)
    {
        var tryCount = IncrementEventTryCount(eventData);
        if (tryCount <= MaxEventTryCount)
        {
            Logger.LogWarning($"Could not perform tenant created event. Re-queueing the operation. TenantId = {eventData.Id}, TenantName = {eventData.Name}.");
            Logger.LogException(exception, LogLevel.Warning);
    
            await Task.Delay(RandomHelper.GetRandom(5000, 15000));
            await DistributedEventBus.PublishAsync(eventData);
        }
        else
        {
            Logger.LogError($"Could not perform tenant created event. Canceling the operation. TenantId = {eventData.Id}, TenantName = {eventData.Name}.");
            Logger.LogException(exception);
        }
    }
    

    Thanks,

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
  • User Avatar
    0
    laura created

    Hi, but when I create a new tenant, 50% of the time it doesn't create the admin user. I have event management with In/Outbox, and even though the database is single, I have different DbContexts for the various services: SaaS, Identity, Administration, etc

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
  • User Avatar
    0
    maliming created
    Support Team Fullstack Developer

    hi

    Do your logs contain Could not perform tenant created event. Re-queueing the operation. TenantId messages?

    Let's make sure the retry mechanism is working.

    Thanks

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
  • User Avatar
    0
    laura created

    Hi,

    no error message in the log, the identy service logs:

    [10:10:56 INF] Sending HTTP request GET http://localhost/health-status [10:10:56 INF] Request starting HTTP/1.1 GET http://localhost/health-status - null null [10:10:56 INF] Executing endpoint 'Health checks' [10:10:56 INF] Executed endpoint 'Health checks' [10:10:56 INF] Received HTTP response headers after 3.881ms - 200 [10:10:56 INF] End processing HTTP request after 3.9598ms - 200 [10:10:56 INF] Request finished HTTP/1.1 GET http://localhost/health-status - 200 null application/json 3.9941ms [10:11:05 INF] Found 1 events in the outbox. [10:11:05 INF] Seeding Permission data... [10:11:05 INF] Sent 1 events to message broker [10:11:05 INF] Seeding admin permissions. [10:11:05 INF] Seeding Identity data... [10:11:06 INF] Start processing HTTP request GET http://localhost/health-status [10:11:06 INF] Sending HTTP request GET http://localhost/health-status [10:11:06 INF] Request starting HTTP/1.1 GET http://localhost/health-status - null null [10:11:06 INF] Executing endpoint 'Health checks' [10:11:06 INF] Executed endpoint 'Health checks' [10:11:06 INF] Request finished HTTP/1.1 GET http://localhost/health-status - 200 null application/json 2.3963ms [10:11:06 INF] Received HTTP response headers after 2.7722ms - 200 [10:11:06 INF] End processing HTTP request after 2.8461ms - 200 [10:11:16 INF] Start processing HTTP request GET http://localhost/health-status [10:11:16 INF] Sending HTTP request GET http://localhost/health-status [10:11:16 INF] Request starting HTTP/1.1 GET http://localhost/health-status - null null [10:11:16 INF] Executing endpoint 'Health checks' [10:11:16 INF] Executed endpoint 'Health checks'

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
  • User Avatar
    0
    AI-Bot created
    Support Team ABP Agent
    • In the microservice template, it’s expected that multiple IDistributedEventHandler<TenantCreatedEto> handlers run in parallel in each service (SaaS/Identity/Administration/etc.). In the Identity service specifically, several EF Core module handlers are registered and they all try to run MigrateDatabaseSchemaAsync() around the same time.
    • Those migration handlers are coordinated using a distributed lock. So when 7 handlers start simultaneously, one gets the lock, the others can fail/timeout acquiring the lock.
    • Your “50% admin user not created” symptom happens when your custom TenantCreatedEto handler (the one that seeds the admin) wraps migration + seeding in a single try/catch and treats a migration lock failure as a fatal failure for the whole operation, but then does not rethrow (so the inbox marks the message as processed and no retry occurs).

    Apply this fix pattern in your custom Identity tenant-created handler:

    1. Do not block seeding on schema migration for shared-db tenants
      For shared database tenants, schema migration is unnecessary. So make seeding independent from migration outcome. Concretely:

      • Run migration in its own try/catch.
      • Always attempt admin seeding in a separate try/catch (or at least outside the migration try/catch).
      • If seeding fails, rethrow so the inbox/outbox retry mechanism can retry.
    2. Only migrate when the tenant has a dedicated connection string
      In the shared-db case, skip MigrateDatabaseSchemaAsync() entirely. (In ABP’s default approach, migration is only meaningful when the tenant has its own database/connection string.)

    3. Ensure failures are not swallowed
      If you override error handling and swallow exceptions, the event will be marked processed. If you want ABP’s retry behavior, you must let the failure propagate (or explicitly re-publish like ABP does internally).

    Practical outcome: even if other module handlers contend on the distributed lock, your admin seeding will still run for shared-db tenants, and if it ever fails it will be retried instead of being silently “processed”.

    Links (verified):

    • EF Core migration event handler base (shows where the lock/retry logic is implemented):
      https://github.com/abpframework/abp/blob/HEAD/framework/src/Volo.Abp.EntityFrameworkCore/Volo/Abp/EntityFrameworkCore/Migrations/EfCoreDatabaseMigrationEventHandlerBase.cs
    • ABP KB about tenant admin user creation/seeding expectations:
      https://abp.io/docs/latest/kb/can-not-login-with-admin-user
    • Background reading on shared vs separate tenant DB behavior (explains why shared-db tenants don’t need per-tenant migrations):
      https://abp.io/community/articles/51nvl4u9

    This response was generated by AI. If it didn’t solve your problem, feel free to post a new answer — our support team will jump in to help you soon.

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
  • User Avatar
    0
    maliming created
    Support Team Fullstack Developer

    hi

    . If the distributed lock times out during migration, the entire block fails

    In this case, the HandleErrorTenantCreatedAsync method will output messages.

    Could not perform tenant created event. Re-queueing the operation. TenantId

    Are you sure it's caused by the distributed lock?

    Thanks.

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
  • User Avatar
    0
    laura created

    I'm not sure. I noted that when I create a tenant in Saas module sometimes the admin user is not created

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
  • User Avatar
    0
    maliming created
    Support Team Fullstack Developer

    hi

    Can you enable debug logging and share the logs.txt when the problem occurs?

    https://abp.io/support/questions/8622/How-to-enable-Debug-logs-for-troubleshoot-problems

    Thanks

    Markdown supported.
    Copy, paste, or drag & drop images and files (max 100 MB per file, 100 MB total per post)
Boost Your Development
ABP Live Training
Packages
See Trainings
Mastering ABP Framework Book
The Official Guide
Mastering
ABP Framework
Learn More
Mastering ABP Framework Book
Made with ❤️ on ABP v10.8.0-preview. Updated on September 28, 2026, 11:44
1
ABP Assistant
🔐 You need to be logged in to use the chatbot. Please log in first.