openssl

Author	SHA1	Message	Date
Patrick Steuer	af1d638730	s390x assembly pack: remove capability double-checking. An instruction's QUERY function is executed at initialization, iff the required MSA level is installed. Therefore, it is sufficient to check the bits returned by the QUERY functions. The MSA level does not have to be checked at every function call. crypto/aes/asm/aes-s390x.pl: The AES key schedule must be computed if the required KM or KMC function codes are not available. Formally, the availability of a KMC function code does not imply the availability of the corresponding KM function code. Signed-off-by: Patrick Steuer <patrick.steuer@de.ibm.com> Reviewed-by: Andy Polyakov <appro@openssl.org> Reviewed-by: Rich Salz <rsalz@openssl.org> (Merged from https://github.com/openssl/openssl/pull/4501)	2017-10-17 21:55:33 +02:00
Rich Salz	e3713c365c	Remove email addresses from source code. Names were not removed. Some comments were updated. Replace Andy's address with openssl.org Reviewed-by: Andy Polyakov <appro@openssl.org> Reviewed-by: Paul Dale <paul.dale@oracle.com> (Merged from https://github.com/openssl/openssl/pull/4516)	2017-10-13 10:06:59 -04:00
Andy Polyakov	236dd46339	sha/asm/keccak1600-armv8.pl: fix return value buglet and ... ... script data load. On related note an attempt was made to merge rotations with logical operations. I mean as we know, ARM ISA has merged rotate-n-logical instructions which can be used here. And they were used to improve keccak1600-armv4 performance. But not here. Even though this approach resulted in improvement on Cortex-A53 proportional to reduction of amount of instructions, ~8%, it didn't exactly worked out on non-Cortex cores. Presumably because they break merged instructions to separate μ-ops, which results in higher operations count. X-Gene and Denver went ~20% slower and Apple A7 - 40%. The optimization was therefore dismissed. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-09-09 19:09:36 +02:00
Rich Salz	b2db9c18b2	MSC_VER <= 1200 isn't supported; remove dead code VisualStudio 6 and earlier aren't supported. Reviewed-by: Andy Polyakov <appro@openssl.org> (Merged from https://github.com/openssl/openssl/pull/4263)	2017-08-27 11:35:39 -04:00
Andy Polyakov	e0584e96c1	sha/asm/keccak1600-armv4.pl: optimize for Thumb-2. Reduce per-round instruction count in Thumb-2 case by 16%. This is achieved by folding ldr/str pairs to their double-word counterparts. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-08-16 20:25:20 +02:00
Andy Polyakov	3c1a60e56f	sha/asm/keccak1600-avx512.pl: fix buglet in SHA3_squeeze tail. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-08-12 12:23:31 +02:00
Andy Polyakov	d9ca12cbf6	sha/asm/keccak1600-armv4.pl: improve non-NEON performance by ~10%. This is achieved mostly by ~10% reduction of amount of instructions per round thanks to a) switch to KECCAK_2X variant; b) merge of almost 1/2 rotations with logical instructions. Performance is improved on all observed processors except on Cortex-A15. This is because it's capable of exploiting more parallelism and can execute original code for same amount of time. Reviewed-by: Rich Salz <rsalz@openssl.org> Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de> (Merged from https://github.com/openssl/openssl/pull/4057)	2017-08-02 23:22:28 +02:00
Andy Polyakov	5d010e3f10	sha/keccak1600.c: choose more sensible default parameters. "More" refers to the fact that we make active BIT_INTERLEAVE choice in some specific cases. Update commentary correspondingly. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-08-01 22:42:35 +02:00
Xiaoyin Liu	bac5b39c96	Fix typo in sha1-thumb.pl Reviewed-by: Tim Hudson <tjh@openssl.org> Reviewed-by: Rich Salz <rsalz@openssl.org> (Merged from https://github.com/openssl/openssl/pull/4056)	2017-07-30 21:26:38 -04:00
Andy Polyakov	c363ce55f2	sha/keccak1600.c: build and make it work with strict warnings. Reviewed-by: Paul Dale <paul.dale@oracle.com> Reviewed-by: Richard Levitte <levitte@openssl.org> (Merged from https://github.com/openssl/openssl/pull/3943)	2017-07-25 21:38:48 +02:00
Andy Polyakov	e3c79f0f19	sha/asm/keccak1600-avx512.pl: improve performance by 17%. Improvement is result of combination of data layout ideas from Keccak Code Package and initial version of this module. Hardware used for benchmarking courtesy of Atos, experiments run by Romain Dolbeau <romain.dolbeau@atos.net>. Kudos! Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de> Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-07-24 21:23:01 +02:00
Andy Polyakov	0d7903f83f	sha/asm/keccak1600-avx512.pl: absorb bug-fix and minor optimization. Hardware used for benchmarking courtesy of Atos, experiments run by Romain Dolbeau <romain.dolbeau@atos.net>. Kudos! Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-07-21 14:12:14 +02:00
Andy Polyakov	64d92d7498	x86_64 assembly pack: "optimize" for Knights Landing, add AVX-512 results. "Optimize" is in quotes because it's rather a "salvage operation" for now. Idea is to identify processor capability flags that drive Knights Landing to suboptimial code paths and mask them. Two flags were identified, XSAVE and ADCX/ADOX. Former affects choice of AES-NI code path specific for Silvermont (Knights Landing is of Silvermont "ancestry"). And 64-bit ADCX/ADOX instructions are effectively mishandled at decode time. In both cases we are looking at ~2x improvement. AVX-512 results cover even Skylake-X :-) Hardware used for benchmarking courtesy of Atos, experiments run by Romain Dolbeau <romain.dolbeau@atos.net>. Kudos! Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-07-21 14:07:32 +02:00
Andy Polyakov	d212b98b36	sha/asm/keccak1600-avx2.pl: optimized remodelled version. New register usage pattern allows to achieve sligtly better performance. Not as much as I hoped for. Performance is believed to be limited by irreconcilable write-back conflicts, rather than lack of computational resources or data dependencies. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-07-15 23:04:38 +02:00
Andy Polyakov	91dbdc63bd	sha/asm/keccak1600-avx2.pl: remodel register usage. This gives much more freedom to rearrange instructions. This is unoptimized version, provided for reference. Basically you need to compare it to initial `29724d0e15` to figure out the key difference. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-07-15 23:04:18 +02:00
Andy Polyakov	c7c7a8e601	Optimize sha/asm/keccak1600-avx2.pl. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-07-10 10:16:42 +02:00
Andy Polyakov	29724d0e15	Add sha/asm/keccak1600-avx2.pl. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-07-10 10:16:31 +02:00
Andy Polyakov	313fa47fea	Add sha/asm/keccak1600-avx512.pl. Reviewed-by: Rich Salz <rsalz@openssl.org> Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de> (Merged from https://github.com/openssl/openssl/pull/3861)	2017-07-07 10:04:33 +02:00
Andy Polyakov	b4f2a462b7	sha/keccak1600.c: internalize KeccakF1600 and simplify SHA3_absorb. Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de>	2017-07-03 18:18:10 +02:00
Andy Polyakov	edbc681d22	sha/asm/keccak1600-x86_64.pl: close gap with Keccak Code Package. [Also typo and readability fixes. Ryzen result is added.] Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de>	2017-07-03 18:18:02 +02:00
Andy Polyakov	b547aba954	sha/asm/keccak1600-s390x.pl: typo and readability, minor size optimization. Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de>	2017-07-03 18:17:55 +02:00
Andy Polyakov	54f8f9a1ed	x86_64 assembly pack: fill some blanks in Ryzen results. Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de>	2017-07-03 18:17:00 +02:00
Andy Polyakov	7807267bed	Add sha/asm/keccak1600-s390x.pl. Reviewed-by: Richard Levitte <levitte@openssl.org>	2017-06-29 21:16:02 +02:00
Andy Polyakov	d6f0c94a65	sha/asm/keccak1600-x86_64.pl: add CFI directives. Reviewed-by: Richard Levitte <levitte@openssl.org>	2017-06-29 21:15:56 +02:00
Andy Polyakov	a1613840dd	sha/asm/keccak1600-x86_64.pl: optimize by re-ordering instructions. Reviewed-by: Richard Levitte <levitte@openssl.org>	2017-06-29 21:15:51 +02:00
Andy Polyakov	a078d9dfa9	sha/asm/keccak1600-x86_64.pl: remove redundant moves. Reviewed-by: Richard Levitte <levitte@openssl.org>	2017-06-29 21:15:45 +02:00
Andy Polyakov	64aef3f53d	Add sha/asm/keccak1600-x86_64.pl. Reviewed-by: Richard Levitte <levitte@openssl.org>	2017-06-29 21:15:09 +02:00
Andy Polyakov	a163e60d95	sha/asm/keccak1600-mmx.pl: optimize for Atom and add comparison data. Curiously enough out-of-order Silvermont benefited most from optimization, 33%. [Originally mentioned "anomaly" turned to be misreported frequency scaling problem. Correct results were collected under older kernel.] Reviewed-by: Rich Salz <rsalz@openssl.org> Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de> (Merged from https://github.com/openssl/openssl/pull/3739)	2017-06-24 09:42:14 +02:00
Andy Polyakov	415248e1e1	Add sha/asm/keccak1600-mmx.pl, x86 MMX module. Reviewed-by: Rich Salz <rsalz@openssl.org> Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de> (Merged from https://github.com/openssl/openssl/pull/3739)	2017-06-24 09:42:08 +02:00
Andy Polyakov	b5cdec2fea	sha/asm/sha512p8-ppc.pl: add POWER8 performance data. [skip ci] Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de> Reviewed-by: Rich Salz <rsalz@openssl.org> (Merged from https://github.com/openssl/openssl/pull/3705)	2017-06-21 16:26:59 +02:00
Andy Polyakov	53ddf7dd05	Add Keccak-1600 modules for PPC64 and POWER8. [skip ci] Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de> Reviewed-by: Rich Salz <rsalz@openssl.org> (Merged from https://github.com/openssl/openssl/pull/3705)	2017-06-21 16:24:36 +02:00
Andy Polyakov	1d23bbccd3	Add sha/asm/keccak1600-c64x.pl [skip ci] Reviewed-by: Bernd Edlinger <bernd.edlinger@hotmail.de> (Merged from https://github.com/openssl/openssl/pull/3708)	2017-06-21 15:21:47 +02:00
Andy Polyakov	5eb2dd88b3	Add sha/asm/keccak1600-armv8.pl. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-15 21:53:30 +02:00
Andy Polyakov	6dad1efef7	sha/asm/keccak1600-armv4.pl: switch to more efficient bit interleaving algorithm. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-08 20:21:31 +02:00
Andy Polyakov	13603583b3	sha/keccak1600.c: switch to more efficient bit interleaving algorithm. [Also bypass sizeof(void *) == 8 check on some platforms.] Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-08 20:21:04 +02:00
Andy Polyakov	367c552790	sha/asm/keccak1600-armv4.pl: add NEON code path. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-06 19:54:29 +02:00
Andy Polyakov	56676f877d	sha/asm/keccak1600-armv4.pl: add SHA3_absorb and SHA3_squeeze. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-06 19:54:24 +02:00
Andy Polyakov	5371810714	sha/asm/keccak1600-armv4.pl: optimization based on profiler feedback. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-06 19:54:19 +02:00
Andy Polyakov	aabfd32910	Add sha/asm/keccak1600-armv4.pl. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-06 19:54:12 +02:00
Andy Polyakov	71dd3b6464	sha/keccak1600.c: add #ifdef KECCAK1600_ASM. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-05 19:35:41 +02:00
Andy Polyakov	22f9fa6e06	sha/keccak1600.c: reduce temporary storage utilization even futher. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-05 19:35:30 +02:00
Andy Polyakov	1ded2dd3ee	sha/keccak1600.c: add another 1x variant. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-05 19:35:07 +02:00
Andy Polyakov	c83a4db521	sha/keccak1600.c: add ARM-specific "reference" tweaks. Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-06-05 19:34:48 +02:00
Andy Polyakov	1f2aff257d	sha/keccak1600.c: implement lane complementing transform ...as discussed in section 2.2 of "Keccak implementation overview". [skip ci] Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-05-30 19:53:14 +02:00
Andy Polyakov	0dd0be9408	sha/keccak1600.c: implement bit interleaving optimization. This targets 32-bit processors and is discussed in section 2.1 of "Keccak implementation overview". Reviewed-by: Rich Salz <rsalz@openssl.org>	2017-05-30 19:52:50 +02:00
David Benjamin	e195c8a256	Remove filename argument to x86 asm_init. The assembler already knows the actual path to the generated file and, in other perlasm architectures, is left to manage debug symbols itself. Notably, in OpenSSL 1.1.x's new build system, which allows a separate build directory, converting .pl to .s as the scripts currently do result in the wrong paths. This also avoids inconsistencies from some of the files using $0 and some passing in the filename. Reviewed-by: Richard Levitte <levitte@openssl.org> Reviewed-by: Andy Polyakov <appro@openssl.org> (Merged from https://github.com/openssl/openssl/pull/3431)	2017-05-11 17:00:23 -04:00
Richard Levitte	74a011ebb5	Cleanup - use e_os2.h rather than stdint.h Not exactly everywhere, but in those source files where stdint.h is included conditionally, or where it will be eventually Reviewed-by: Rich Salz <rsalz@openssl.org> (Merged from https://github.com/openssl/openssl/pull/3447)	2017-05-11 21:52:37 +02:00
Andy Polyakov	ce1932f25f	sha/sha512.c: fix formatting. Reviewed-by: Richard Levitte <levitte@openssl.org>	2017-05-05 17:04:09 +02:00
FdaSilvaYY	69687aa829	More typo fixes Fix some comments too [skip ci] Reviewed-by: Tim Hudson <tjh@openssl.org> Reviewed-by: Richard Levitte <levitte@openssl.org> (Merged from https://github.com/openssl/openssl/pull/3069)	2017-03-29 07:14:29 +02:00
Andy Polyakov	6cbfd94d08	x86_64 assembly pack: add some Ryzen performance results. Reviewed-by: Tim Hudson <tjh@openssl.org>	2017-03-22 10:58:01 +01:00

1 2 3 4 5 ...

523 commits