Unfortunately I cannot copy the graphs to this post, but this is the response I have had from VisualWare, who sell their diagnostic software to companies who need to verify their connection problems to prove their point to ISP's.
The guy there is trying to establish if he can find the problem before I purchase the software. I have sent him one test result (TCP stack test) and this is what he thinks so far: (sorry for the long post!)
Hi Neil
This looks like it is a fairly good test, the main issue is the red spikes that align to blue spikes are packet loss events, either lost altogether or what is called a TCP black hole. The blue spike is the delay caused by the event. They don�t look like much (and this was a good test I think) they do have an impact. 30ms on a 3Mbs connection is losing a throughput of nearly 100Kb. However that said this looks like a relatively clean test. You indicated you were getting TCP delay much higher than this. Is that correct
Don�t get confused with QoS, it is a measure of consistency. It is very important as it relates to the flowing of data evenly but it is not a speed related measure. Yours is high because the breakdown is symmetrical.
[0,1600) = 541,173 bytes (2,705,865 bps)
[1600,3200) = 539,956 bytes (2,699,780 bps)
[3200,4800) = 537,719 bytes (2,688,595 bps)
[4800,6400) = 537,328 bytes (2,686,640 bps)
[6400,8001) = 538,983 bytes (2,693,231 bps)
The modal pause analysis shows that your data is flowing at about 3Mbps (excluding the delays) is that correct
Pause frequencies
4ms | 917 | #################
5ms | 583 | ###########
3ms | 159 | ###
6ms | 76 | #
33ms | 1 |
1ms | 8 |
8ms | 5 |
2ms | 13 |
12ms | 1 |
10ms | 2 |
0ms | 5 |
15ms | 2 |
7ms | 5 |
31ms | 1 |
11ms | 2 |
32ms | 1 |
28ms | 2 |
17ms | 1 |
9ms | 3 |
20ms | 1 |
18ms | 1 |
24ms | 1 |
26ms | 1 |
Preliminary guess at modal pause = 4 ms
Considering 870 pauses of 4 ms
Considering 555 pauses of 5 ms
Considering 145 pauses of 3 ms
Modal pause = 6690 / 1570 = 4.2
Bandwidth = 342630 --> 340000
Ideally I need to see a bad test, my guess is that you suffer from what is called a congestion pattern. Just like the M25 at certain times oversubscribed networks slow down because of too much traffic. If I can get a test log of a bad result that will help.
UPLOAD is showing a fairly consistent service however there is one very HUGE interrupt loss event or 1000ms. That will impact almost any application. It is a whole second without data. Also if you review the data excluding the spike to get a better scale you can see other delays. When you get delays far greater then the trip latency (which is 37ms) then it will almost certainly trigger duplicates. Unfortunately duplicates consume bandwidth for not value to the application as the packets are duplicates and are just thrown away.
A capacity test will not help to solve this problem, an application test is far more useful as it tracks the data from an application standpoint. However if you want me to investigate why this is not working just let me know. The feature is not available in PCLite however.
If you want to do a 30 hour automated congestion test let me know and I will set up the download for you. Otherwise in the meantime I will wait to see a bad result report. Can you reconfirm your contracted service speeds for me. I want to be sure the pattern is valid for the service. What it is showing is 3Mbps currently as a flow rate.
Best wishes
Julian
In broadband terms "up to" is a pseudonym for "we don't care".