Repository navigation
Failed test: test-child-process-fork-regr-gh-2847 #3635
Description
Activity
- addedwindowsIssues and PRs related to the Windows platform.Issues and PRs related to the Windows platform.
on Nov 2, 2015 - addedchild_processIssues and PRs related to the child_process subsystem.Issues and PRs related to the child_process subsystem.
on Nov 3, 2015 Also failed in #3631 (comment), I'm not sure if that is one of those two runs.
The test was added in 36b969f by @indutny and then modified in e9e837c by @mhdawson.
@gireeshpunathil can you take a look at the windows failures since you worked on the the most recent change.
@mhdawson , sure, I will debug to see what the issue is.
When run in many iterations, I am able to see it failing with the original test case as well (36b969f). So this is not a regression.
There are three types of errors observed in total. With e9e837c (a) and (b) seem to have eliminated, but (c) persists.
(a)
events.js:141 throw er; // Unhandled 'error' event ^ Error: connect ECONNREFUSED 127.0.0.1:12346 at Object.exports._errnoException (util.js:915:11) at exports._exceptionWithHostPort (util.js:938:20) at TCPConnectWrap.afterConnect [as oncomplete] (net.js:1065:14)(b)
events.js:141 throw er; // Unhandled 'error' event ^ Error: write EMFILE at exports._errnoException (util.js:915:11) at ChildProcess.target._send (internal/child_process.js:606:18) at internal/child_process.js:469:16 at Array.forEach (native) at ChildProcess.<anonymous> (internal/child_process.js:468:13) at emitTwo (events.js:92:20) at ChildProcess.emit (events.js:172:7) at handleMessage (internal/child_process.js:686:10) at Pipe.channel.onread (internal/child_process.js:440:11)(c)
events.js:141 throw er; // Unhandled 'error' event Error: read ECONNRESET at exports._errnoException (util.js:915:11) at TCP.onread (net.js:544:26)The summary of the test case is to validate that the worker process being closed have its _channel field nullified. To re-inforce this assertion, a number of requests are sent from the master. The problem with the test case is that it does not handle the failure scenarios which can occur when the server is shutdown at arbitrary time intervals:
i) When an nth connection request is issued, the (n-1)th request could have resulted in a worker closure, followed by server shutdown. This results in failure (a)
ii) When an nth send request is issued, the (n-1)th request could have resulted in a worker closure, followed by server shutdown. This results in failure (b)
iii) Even when (i) and (ii) are false, the server socket could be closed in between several TCP protocol message transfers which constitute the connect or send. This results in failure (c).Tweaking test case further, I see that even sending just two messages to the worker is sufficient to cause (c). More interestingly, as the error type (c) is coming from the TCP stack, detached from the net connect, send APIs, their error callbacks are incapable of catching this error. This essentially means that we cannot reliably pass messages, or suppress the errors.
Any suggestions to improve the test case is welcome.
- added 6 commits that reference this issue
on Nov 24, 2015 - added 2 commits that reference this issue
on Dec 17, 2015 46 remaining items
- added a commit that references this issue
on May 19, 2017 - added a commit that references this issue
on Jul 17, 2017
Looks like this test is starting to fail in our windows environment: fail 1, fail 2