Thank you Thomas for your responses.  I'll do some more digging on my
side.   Appreciate all of your contributions to this project!!

- Rob


On Wed, Sep 16, 2026, 7:44 AM Thomas Eckardt <[email protected]>
wrote:

> If 'ConnectionLog' and/or 'SessionLog' are set to verbose or higher, I
> expect the following error shown in the maillog, if assp is unable to close
> a tcp-socket to 'smtpDestination'.
>
> error: unable to close Socket SOCKET - PERLERROR - SYSTEMERROR
>
>
>
> Notice:
> - the assp development version you use is 18 builds behind the current
> build - ASSP 2.8.2(26253)
>
> - Perl 5.26.1 status:
> Release Date: September 2017
> Active Support Ended: May 2019
> Security Support Ended: May 2020
> Current Status: Unsupported and obsolete
>
> - ASSP 2.8.2(24261) is used by only 3 instances (including yours)
> worldwide - so I don't expect any feedback for this build
>
>
> *1. Is there a known issue with smtpDestination connection
> handling/closing*
> *   in this version or fixed in a later release?*
>
> please read the changelog.txt by your self! so I could skip the copy and
> paste.
>
> build 24261 2024-09-17
> fixed:
>
>  - related to https://assp.thockar.com/forum/viewtopic.php?t=3700
>    If enableINET6 was used, binding to a listener after an assp restart
> failed on some systems, because assp used the
>    deprecated IO::Socket::INET option 'Reuse' also in calls to
> IO::Socket::IP, where this option is not supported.
>    Now assp uses the 'ReuseAddr' option instead, which is supported in
> both modules.
>
> added:
>
> - related to https://assp.thockar.com/forum/viewtopic.php?t=3700
>   If $disable_SO_REUSEPORT is set to 0 or 2 assp tries to use the
> SO_REUSEPORT socket option
>
> our $disable_SO_REUSEPORT = 1;           # (0/1/2) disable the
> SO_REUSEPORT socket option - there is no assp version which ever used this
> option
>                                          # notice: windows never has this
> socket option - so leave this value at 1
>                                          #  0 - do not disable - try it,
> but if not supported by the OS it is not used and a load warning is
> produced for the module 'Socket'
>                                          #  1 - disable this socket option
> and do not try to use it
>                                          #  2 - do not disable - try it,
> but if not supported by the OS it is not used and silently ignored
>
>   You may try to play around with this value, if you get unexpected errors
> for assp listeners while bind or reuse.
>
> !!! the related code is currently untouched since build 24261
>
> *2. Is there a config option to disable backend connection reuse/pooling
> (if*
> *   that's what's happening here), or to force ASSP to close backend*
> *   connections immediately after each transaction rather than potentially*
> *   holding them open for reuse?*
>
> see above
>
> ASSP needs and tries to close every connection. REUSE.... options are used
> by assp at creation time of a socket listener, and should be handled by the
> underlying system components (kernel, libc ...)
> ASSP uses the socket it gets from the system at socket creation time. It
> uses REUSEADDR for every (but only) listener - usage of REUSEPORT depends
> on config ($disable_SO_REUSEPORT) and system support.
>
> Connections to the 'smtpDestination' backend are *tcp client connections*,
> and should be not affected by any of these parameters.
>
> I just had a look in to my "small" knowledgebase and found a nice
> explanation of a very special tcp CLOSE_WAIT case:
>
> "The connection might also go to CLOSE_WAIT even if the server didn't yet
> invoke accept() on the socket and the client already close()"
>
> So - if your backenserver gets the connection, but does'nt do a
> sock->accept() for any reason, it may happen that the connection stays in
> CLOSE_WAIT at the assp system - even assp closed the socket clearly!
> It is currently unclear to me what's going on in this case (e.g. the assp
> system sends a RST instead of a FIN ?? - the FIN to the backend is the
> missing part here)
>
> If this (or anything similar) happens, you'll not see the assp error
> message in the log
> "error: unable to close Socket SOCKET - PERLERROR - SYSTEMERROR"
> because I expect the system returns success on the assp close() call.
>
> It makes sense to me, that SSL connections are not affected. SSL
> connections are mostly completely and hard closed at both sites - for
> security reasons.
>
> btw: did you change/update/install anything on any of the involved systems
> ?
>
> This is just a shot into the dark. Checking both, the backend log and assp
> log (time based) using some increased log levels may lead into some more
> brightness.
>
> I'm sure, I can't help you any further. Possibly someone else is able to
> bring some light into this.
>
> *3. Has anyone else seen CLOSE_WAIT accumulation specifically on the*
> *   smtpDestination leg rather than the client-facing leg?*
>
> there is  no related issue reported for build 24261 and all higher builds
>
> Thomas
>
>
>
>
> Von:        "Robert Ellsworth" <[email protected]>
> An:        "For Users of ASSP" <[email protected]>
> Datum:        15.09.2026 16:00
> Betreff:        [Assp-user] [Bug Report] Persistent CLOSE_WAIT socket
> leak on smtpDestination (backend) connections - ASSP 2.8.2(24261)
> ------------------------------
>
>
>
>
>
> Subject: [Bug Report] Persistent CLOSE_WAIT socket leak on smtpDestination
> (backend) connections - ASSP 2.8.2(24261)
>
> Hi all,
>
> I'm seeing a reproducible file descriptor leak on our ASSP installation
> that
> eventually leads to "Too many open files" errors and worker crashes if left
> unaddressed. I've isolated it fairly precisely and wanted to share the
> evidence in case it's a known issue or others are hitting it silently.
>
> ## Environment
>
> - ASSP version: 2.8.2(24261)
> - OS: Debian/Ubuntu (Linux), Perl 5.26.1
> - Run mode: systemd-managed daemon (Type=forking, PIDFile), ulimit raised
> to
>   65536 (soft and hard) to rule out an undersized limit as the cause
> - smtpDestination:=192.168.1.54:25|SSL:*192.168.1.54:465*
> <http://192.168.1.54:465/>
>   (backend is Zimbra, both plaintext :25 and SSL :465 destinations
> configured)
>
> ## Symptom
>
> Open file descriptor count on the *assp.pl* <http://assp.pl/> process
> climbs steadily over
> several hours, with no plateau, until it approaches the ulimit ceiling and
> worker threads start failing with:
>
>     Worker_X accept to client failed IO::Socket::INET=GLOB(...) (timeout:
> 2 s) : Too many open files
>     main exception: error: AsspSelfLoader is unable to load code from file
> .../sl-cache/main-resetFH.sl - Too many open files
>
> This has now happened on multiple separate occasions (Aug 28, Sep 1, Sep 4,
> and again this morning, Sep 15), each following the same growth pattern:
> flat baseline FD count (~40-50) for hours, then a sustained climb beginning
> at an arbitrary point in the day and continuing until intervention.
>
> ## Isolating the leak
>
> Using `lsof -p <pid>`, I confirmed the growth is entirely in socket file
> descriptors, not regular files, disk cache handles, or DNS-related opens
> (REG/CHR/DIR counts stay flat throughout):
>
>     12:13:47   86 REG   46 IPv4   7 CHR   5 sock   2 DIR
>     12:28:47   86 REG  115 IPv4   7 CHR  11 sock   2 DIR
>     12:43:48   87 REG  293 IPv4   7 CHR  50 sock   2 DIR
>     12:58:48   86 REG  361 IPv4  92 sock   7 CHR   2 DIR
>     13:13:49   88 REG  536 IPv4 126 sock   7 CHR   2 DIR
>     13:43:56   86 REG  668 IPv4 255 sock   7 CHR   2 DIR
>
> Breaking down socket state with `lsof -p <pid> -i` during an active leak
> event this morning:
>
>     335 CLOSE_WAIT
>      28 ESTABLISHED
>      12 LISTEN
>
> CLOSE_WAIT dominates - meaning the remote peer has already closed its end
> (sent FIN) but ASSP has not called close() on its side.
>
> Breaking the CLOSE_WAIT sockets down by remote host:
>
>     246 our.backend.mailserver   <- our smtpDestination backend
>      72 *mailsrv.graphic.com.gh* <http://mailsrv.graphic.com.gh/>
>       2 *mail-qv2-f12.google.com* <http://mail-qv2-f12.google.com/>
>       1 (various other single-connection external senders - normal
> background noise)
>
> The overwhelming majority (73%+) are connections to our own backend/
> destination server (Zimbra), not inbound client connections. Inspecting the
> port on these:
>
>     perl 7243 root 46u IPv4 ... TCP 
> assp.eomega.org:37038->our.backend.mailserver:smtp
> (CLOSE_WAIT)
>     perl 7243 root 51u IPv4 ... TCP 
> assp.eomega.org:43076->our.backend.mailserver:smtp
> (CLOSE_WAIT)
>     ... (continues, ~295 similar entries)
>
> Nearly all of these are the plaintext `smtp` (port 25) destination, not the
> SSL (:465) destination - so this does not appear to be an SSL-shutdown
> issue, but rather plain-TCP backend connections that are not being closed
> after the delivery/proxy transaction completes.
>
> ## What this looks like to me
>
> It appears that when ASSP proxies a message through to smtpDestination and
> the backend (Zimbra) closes its end of the connection after completing the
> SMTP transaction, ASSP does not follow up with its own close() on that
> socket - leaving it in CLOSE_WAIT indefinitely. Over hours, under normal
> mail volume, this accumulates into the thousands and exhausts the process's
> file descriptor limit.
>
> ## What I've ruled out
>
> - Undersized ulimit: raised to 65536, leak still occurs, just takes longer
>   to become fatal
> - DNS resolver file-open failures: present in logs but unrelated; these are
>   a side-effect of FD exhaustion (resolver can't open /etc/protocols), not
>   a cause
> - SSL socket shutdown handling: leak is concentrated on the plaintext :25
>   backend connection, not :465
> - Spam-flood/PenaltyBox tempfail handling: this was our working theory
>   after an earlier incident (leak did correlate with a spam sender getting
>   tempfailed repeatedly), but this most recent capture shows the leak is
>   overwhelmingly backend-connection-related regardless of that theory -
>   possibly a contributing trigger, but not the core mechanism
> - tcp_fin_timeout / kernel-level socket reclaim: not applicable, since
>   CLOSE_WAIT is governed entirely by the application, not any kernel
>   timeout
>
> ## Questions for the list
>
> 1. Is there a known issue with smtpDestination connection handling/closing
>    in this version or fixed in a later release?
> 2. Is there a config option to disable backend connection reuse/pooling (if
>    that's what's happening here), or to force ASSP to close backend
>    connections immediately after each transaction rather than potentially
>    holding them open for reuse?
> 3. Has anyone else seen CLOSE_WAIT accumulation specifically on the
>    smtpDestination leg rather than the client-facing leg?
>
> Happy to provide additional lsof/tcpdump captures, debug logging, or test
> config changes if it helps narrow this down. Currently mitigating with a
> raised ulimit + systemd auto-restart to avoid crashes while we get to the
> bottom of the actual leak.
>
> Thanks,
> Rob_______________________________________________
> Assp-user mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/assp-user
>
>
> _______________________________________________
> Assp-user mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/assp-user
>
_______________________________________________
Assp-user mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/assp-user

Reply via email to