Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I provided exactly one example and that was a comparison between XML and JSON. The follow up was just an additional remark that idiomatic XML would be even shorter.

I am afraid I cannot follow your argument about complexity as any parser worth its money WILL BE ABLE to parse XML without any significant overhead and - as I mentioned numerous times - issues do not stem from how an XML document is built but from the fact that (mostly) Java parsers attempted to implement about each possible OOP pattern instead of going for KISS. The code over-engineering is the issue, not the document structure. I am really not sure what point you are trying to make and why you are diverting from the actual topic I addressed.

I am slightly surprised about your bias remark, as it seems to be you who is strongly biased towards JSON despite my initial comments as well as the "significant benefits" as I never made any such claim, on the contrary it is you once again who seems to make such about JSON. Maybe you can post the "hard data" you referred to.

Bottom line (once again), XML and JSON are extremely similar and saying one is better than the other simply shows lack of experience. Then of course, if you parse 70 kilobyte JSON documents with a lean parser, but parse 12 megabyte XML documents with a typical Java parser, nobody needs to be surprised the latter will perform abysmally compared to the former, but that would be a whole different subject.



To answer your followup post, your defence of XML is interesting but I'm not referring to milliseconds, its about hours. Not talking about bits or bytes, I'm referring to gigabytes of RAM and i/o. Compressed data streams assumed, with data i/o requirements a definite important factor but not the only one.

I'm not an evangelist for JSON, I'm someone who ran tests and came to conclusions with the help of multiple others. These weren't a generic benchmark for random or general academic purposes. These were representative samples of datasets we're actively going to be or actually are already using.

Even in my other interests I'm using JSON for configuration and data transfer. It shines there quite nicely. XML was generally suitable but its verbosity didn't provide any real advantage and the library support tried to drag in too many dependencies. TSV files weren't suitable even though they were simpler and we had control of the data sources.

You mention Java / Javascript but neither is what we're using. There's probably some irony in not using javascript for JSON i/o but it is what it is. (The purists will agree there's no requirement and so do we). You also didn't mention in passing any of the other interchange / file and document formats we actively compared. JSON / XML etc were just some of the candidates.

Thank you for letting me know that the teams I work with demonstrate "lack of experience". I've forgotten which logical fallacy that is but I'll leave that to someone else to know or look up. I'm just glad I've kept beginner's mind: its a key aspect of neuro-plastic mindset. Its a requirement for keeping an open mind.

We won't be ignoring our testing on real data subsets. The results are clear enough.

(The samples we tested with were around 5MB, 50Mb, 1GB, 10GB and 50GB in size as various combinations of lists and trees etc.)


Most definitely no offence meant, but if you are talking about gigabytes in the context of JSON, XML, or any other text based format you are doing something wrong. And yes, in this case I will stand by my "lack of experience" concerning your entire team, I am sorry.

However you havent really addressed your use case anyhow but you just threw keywords around - gigabyte, IO, compressed, etc. You might want to elaborate on where you have to use XML files of the size of 50 gigabytes.

It doesnt really matter what you are using, Java was just one example. If the XML parser you employ has similar issues you must not be surprised if the outcome is similar. And you seem to be coming back over and over to software support (dependencies). Yes, particularly Java was poor when it came to that but as I said quite some time ago, that is an issue with that software not the document format.

I will disregard the 10M+ tests, but could you publish somewhere the results of the 5MB files?

Again, JSON and XML are way too similar to be anywhere close to what you described and aforementioned benchmark highlighted that. Yes, its dataset is average but I am sure you'll be able to extrapolate that for larger sets.

Apart from the apparent improper use for data of that magnitude, I could only imagine you used an XML parser that simply was not fit for the task and if you do that you shouldnt be surprised that it does not work.


Wow. I predict any kind of collaboration would involve a painful set of further interactions with little benefit.

For the benefit of the probably only two others in the studio audience (who are probably currently both facepalming), we tested with multiple libraries, multiple languages, multiple OS and multiple data subsets. We found in our particular experience that JSON worked the best across our criteria using a representative sample of our datasets. Nowhere did I say XML is always the wrong choice for others. I vaguely recall I wrote I'm now none too keen on XML but have used it in the past. For some things I'd actually choose TSV over XML but thats on fairly, hopefully, obvious cases. I think XML's verbosity is actually its strength but that it has tradeoffs which are quite real. This should not come as a revelation to anyone.

I shared a necessarily limited snapshot of an experience I had and an opinion I formed based on it. I think others can do their own testing as I expect they will anyway. They will confirm or deny based on what they are doing. Especially the opposite case of large imports in XML being faster than everything else. That's completely fine by me.

You've definitely made too many assumptions based on too little data. You didn't even ask what industry this was for. Or what kind of data it was. Or even what disparate systems were involved such that we'd end up with something you state are inappropriately large compressed text files. You disregarded the use of "keywords" such as gigabytes or even compression in general as if those should be unimportant to us. Or why we would use JSON at all. Then you make judgements. Fairly condescending ones at that. This shows a general lack of awareness across several aspects of life in general. For the sake of both of those other people still following this chain, I'll finish here. Life is too short.


Ehm, I did not ask? I very much did so

> However you havent really addressed your use case anyhow but you just threw keywords around - gigabyte, IO, compressed, etc. You might want to elaborate on where you have to use XML files of the size of 50 gigabytes.

I even asked if you could provide that one 5 megabyte file. I take your response as you cant.

I really have the feeling we are going in circles here and you seem to want to resort to ridicule at this point, which will make the discussion pointless.

I believe I have made my point very clear from the start, elaborated more than once what my stance on this subject is, and even dug out some benchmarks. If none of that pleases you or makes you understand what I was actually trying to say, then I am terribly sorry but it is pointless.

And I'd appreciate if you could point out where I was "condescending", as I would object to that, except for the "lack of experience" and I still stand by that given the information you have revealed so far.


Not that it actually matters much[1] but as you insist very much on performance here a benchmark

http://www.navioo.com/ajax/examples/json/test.php

In the first example XML actually is a tad faster, in the second example it is practically a tie (JSON wins by three or four milliseconds).

[1] Are we seriously arguing about milliseconds when it comes to document parsing?




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: