Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- Open page, read headline, read first sentence, emdash, close page.by jttnr
- rclone has been the real fully featured ultra power alternative to rsync over the past 4-5 years.
How do you compare Syq to rclone ?
by QuiCasseRien - As for speed, see https://greaber.github.io/syq-bench/rclone.html. As for other stuff, it depends what you use rclone for.by greaber
- Very cool!
How does the direct TCP mode work?
by hikarudo - Currently, there is just a fixed range of ports that it will use if they are open and you don't pass `--no-tcp`. It does its own encryption over the TCP connections. Also worth knowing is that for some connections (e.g. if there are any dropped packets), you may get much better performance if you can enable BBR congestion control. https://greaber.github.io/syq/server-tuning.html#test-conges...by greaber
- Wow, interesting! How is the performance compared to rclone though?by Zenul_Abidin
- I did some benchmarks against rclone today. https://greaber.github.io/syq-bench/rclone.html Syq was sometimes much faster and never more than a few percent slower. I didn't do any tuning of syq since it is meant to perform well without any special options, but I did tune a couple of rclone options. It's possible that more extensive rclone tuning would yield better results though. Please let me know if there is anything I need to try!by greaber
- Why is this better than rsync?by mika6996
- Browsing the repo, I can see a couple of advantages:
- Single contributor. Software bugs are generally caused by developers writing code. By reducing the number of contributors, syq has cleverly reduced the surface area for defects to sneak in.
- Distribution via shell script. rsync is bundled in most linux distributions, which means you have to deal with annoying software updates from time to time. Syq, OTOH, tells you to pipe curl into bash to execute a shell script, which means one-and-done installation and maintenance.
- UI clarity. If you search StackOverflow for rsync, you'll see thousands of questions asking how to accomplish various tasks with rsync, both straightforward and arcane. By contrast, `syq` isn't even a tag on SO. The obvious conclusion here is that rsync's interface is so byzantine, and its documentation so poor, that users must resort to asking strangers for help, a problem obviously not shared by syq.
by nimih - There are a bunch of reasons, some of which are explained in the description and in the docs. For me, the most important benefit is speed of copying. Note that syq doesn't currently implement rsync's delta merge algorithm, which allows it to avoid copying data that is already in the destination file but at a shifted offset. I might implement this (or an enhanced version of it) in the future. If your workloads have a lot of cases like this then syq might not be better than rsync for you, but I found that for me this rarely came up.by greaber
- I'm eager to replace battle-tested rsync with a new half-baked tool that no one uses for a tiny speed improvement. I'm also going to enable automatic updates of that tool too.
- You live up to your username!by xz18r
- "Tiny" speed improvement? According to the benchmarks, it's almost 10x the speed of rsync!
Really serious benchmarks never contain any information about network characteristics, options used or buffer sizes, and these are indeed very serious. The "details" button produces huge bar graphs. No disappointment there.
Absolutely looking forward to 10x my speed, perhaps even transferring gigabytes per second on my gigabit link.
by xorcist - I'm seeing 5x speedups on my real workloads, and you can see some synthetic benchmarks here or run your own https://greaber.github.io/syq-bench/by greaber
- Thank you for showing this. I'll echo what some others say: if you're claiming to be faster than rsync, you should demonstrate why front-and-center. There have been decades of documentation on how rsync optimizes transfers, so it's a lot to live up to.
As an aside, It's perhaps an indictment of our networking landscape that multiple parallel connections between 2 specific machines would accelerate a transfer. I would've expected a single TCP connection to be able to saturate a line. Or perhaps the parallel connections seeks to amortize the per-file setup overhead?
by AceJohnny2 - I entirely agree that multiple connections shouldn't actually be necessary for speed, but they help in multiple situations (long-distance transfers, same DC between servers, NFS), and it's not just about amortizing per-file setup overhead.
There is some info on the optimizations in the docs, but I agree that a more complete technical explanation of all the things syq does could be useful. I will work on one. On the other hand, I also tried hard to make it just go fast without needing the user to understand why it is fast or tune anything. For instance, the number of connections is auto-tuned by default.
by greaber - > I would've expected a single TCP connection to be able to saturate a line
There's several factors at play that make this (usually) not the case.
Besides physical latency (which includes those added by any VPNs/tunnels/etc., some of which may be internal to an ISP along the route and outside of your control), there's other things like the TCP window sizes / window scaling option[1] that can affect single stream performance, and those type of parameters can differ by OS/interface type on both ends.
Also for SSH specifically, it has its own fixed buffer size that also limits throughput unless you're using the HPN-SSH fork[2].
- > I also added cool features like the ability to maintain a persistent ssh connection to the server for fast one-offs
rsync already supports this via OpenSSH's ControlMaster directive. Bonus, it speeds up every connection to that server rather then a single tool's.
> the ability to do direct remote-remote transfers without forwarding your ssh agent (by using restricted ssh keys on the receiver that will only execute a specific request signed with the key on your laptop).
You can control your agent forwarding in your local ssh client configuration. And with modern ProxyJump, you don't need to forward your agent at all regardless of how many bastion hops are between you and your final target.
- Yeah, ControlMaster is cool and was an inspiration, but since the syq client is always talking to the same receiver on the server, we can save a few round trips that ControlMaster still needs, so syq persist feels faster. This architecture is also what enables reverse mode, where you can download to your laptop while working on the server.
ProxyJump doesn't help when you want to copy files from server A to server B without giving server A an agent that can do arbitrary things on server B. There is really no alternative to using a restricted authorized key on server B for this scenario.
by greaber - I thought I'd share a more positive comment. I work on a tool in a similar area (moving stuff around) but on a higher level. I currently use rsync or rclone for the file moving parts. So to me your tool looks interesting and promising, thanks for sharing it :)
Some suggestions for improvement for your website/Github:
- I wanted to understand how it differs from rsync. A comparison page would be helpful.
- A comparison to rclone would also be interesting, as it also uses multiple parallel connections.
- As it target more advanced users I think it would be helpful to provide more advanced technical insights on how your transfers exactly work. Also how are edge cases handled, what tests are done? Moving and copying files is critical, and you need to inspire trust in your project among users.
by hxseven - Thanks! Yeah, I was a little surprised (maybe I shouldn't have been) at how negative some of the comments were. I will think about what additional explanations I can provide that will help people. Regarding rclone, I have also spent a bunch of time looking at faster transfers to and from S3-compatible storage, but that problem is somewhat different, and I didn't look in depth at how rclone is implemented. FWIW, in my experience, s5cmd is generally a bit faster than rclone even if you tune rclone options. (But rclone is more flexible.) My guess is that s5cmd is already close to the performance ceiling, but I don't really know.
As for edge cases, I will also try to document that more. In general, when I'm not sure what semantics to go for, I try to either copy what rsync does or do something safer. For instance, one thing I am looking at now is the best way to handle cases where a source path cannot be represented on the target filesystem or where two source paths would collide (e.g. due to unicode normalization or case insensitivity). Tentatively, my inclination is to try to fail before copying any files when possible. Currently, syq doesn't do this (but neither does rsync).
by greaber